Incident drills your AI runs
Calmdrill is a kit of files you add to the Claude or ChatGPT account you already have. Your AI becomes the Game Master. It runs a realistic outage, breach or continuity drill with your team, then writes the After-Action Report for you to review and sign. No facilitator, no platform, about 45 minutes.
This is an exercise message.
The first inject of scenario CD-01, as it opened an internal scripted playtest at Harbourline, a fictional company. The page is the scenario’s own. The response and the reply are illustrative, written from the playtest log. The controller is your own Claude or ChatGPT, working from the Calmdrill files.
Paper colour says who a sheet is for. White is for everyone taking part. Canary is for the controller, which is your AI, and players do not see it. Pink is the evaluator’s record, which is what gets filed. Green is the order copy, for whoever holds the budget.
Four ways to run an incident drill
Most teams have an incident plan that has never been rehearsed. Your auditor will ask how you tested it. Your customers assume you have. Your on-call engineer is about to find out at 2 a.m.
Three of the four ask for a budget, a procurement cycle or a colleague who can play Game Master convincingly.
| Option | Cost | Needs | Covers | Evidence |
|---|---|---|---|---|
| Consultant-led tabletop | $5,000–$50,000 per exercise | Weeks of scheduling, an external facilitator | Usually one security scenario | Good report, once a year |
| Enterprise simulation platform | Enterprise pricing, usually quoted on request | Procurement, onboarding, seats | Mostly cyber-security | Platform reports |
| Free PDF pack | Free | Someone who can play Game Master convincingly | Mostly cyber-security, generic | Write it up yourself |
| Calmdrill | $149 once, whole organisation | A Claude or ChatGPT account you already have | Outages, data loss, security and continuity, in 12 scenarios | An After-Action Report drafted for you to sign, every time |
How a drill runs
Ten minutes to set up, about 45 to play and 15 to debrief. The Who column says who acts. Nobody on your team has to prepare or facilitate.
| Elapsed | Activity | Who | What happens |
|---|---|---|---|
| 00:00 | Set up 10 minutes | You | Add the Calmdrill files to a Claude Project or a ChatGPT project, or install the Claude Skill. Say “let’s drill”. Pick a scenario, a mode and who is playing. |
| 00:10 | Play about 45 minutes | Game Master | Keeps a simulated clock, shows you logs and dashboards when you look, fires injects on schedule, and plays the CEO, the angry customer and the lawyer who won’t answer the phone. |
| You | Investigate, decide and keep people informed, as you would at 2 a.m. | ||
| 00:55 | Debrief and file 15 minutes | Game Master | Reveals what really happened, scores six dimensions with evidence from your own decisions, and writes the After-Action Report. |
| You | Agree actions with owners and dates, then review the report and sign it. |
- Three modes
- Team, for three to eight people around a screen. Solo on-call, to practise being paged. Leadership, for decisions about customers, money and regulators.
- Your stack, or ours
- Every scenario runs at Harbourline, a detailed fictional SaaS company. Or describe your own stack in a few lines and the Game Master re-skins the drill to your services and tools.
- Nothing leaves your account
- It runs entirely inside your own AI workspace. We never see your drills, your team or your reports. There is no platform to procure or security-review.
Canary sheet. For the controller. Your AI works from it and players do not see it.
What your AI knows and the players do not
Each scenario has one fixed hidden truth, and the Game Master keeps to it. Look in the right place and you find the real clue. Look in the wrong place and you find normal-looking data, as you would at 2 a.m. Restart everything in a panic and things get worse, as they would in production.
It is told never to rescue you early, never to interpret the evidence for you and never to invent facts that contradict the scenario. A human exercise controller works to the same rules.
| No. | Sim time | Deliver | From | Via | Inject |
|---|---|---|---|---|---|
| 01 | T+00:00 | At start | PagerDuty | Page | [TRIGGERED] auth-svc: login_success_rate < 80% for 5m (currently 71%) |
| 03 | T+00:10 | Always | @BenBrightwater | Social | "Can't log into @Harbourline, got invoices to send before surgery opens. Anyone else? #down" Three replies: "same". |
| 04 | T+00:15 | If no status page update yet | Priya Nair, Head of Support | DM | "14 tickets saying they can't log in. Is this us? What do I say?" |
| 05 | T+00:20 | If the team pages or messages the CTO | Marcus Hale, CTO | Joins | Groggy, asks for a summary. Then wants to "just have a look at the cluster myself". |
| 10 | T+00:45 | Always | Brightwater Dental, ops manager (Enterprise) | Asks for an ETA, and whether their scheduled 08:00 invoice run will go out. | |
| 12 | T+01:00 | If still unresolved | Dana Okafor, CEO | DM | "I need a two-line version for the board chat, and do we owe anyone service credits?" |
The events list is one part of a scenario file. Each file also holds the hidden truth, realistic signals, tempting but risky moves, a four-step hint ladder, scenario-specific scoring and debrief questions.
Pink sheet. The evaluator’s record. What was done and when, and what gets filed.
The report your AI writes
Every drill ends with a structured report built from the exercise log: participants, objectives, timeline, rubric scores with evidence, what went well, what to improve, and owned actions with due dates.
This extract comes from an internal scripted playtest at Harbourline, a fictional company. Its four players were scripted to make the classic mistakes, so it shows what the report says when things go wrong.
The kit also includes an exercise plan template, an attendance record, an improvement tracker and a framework mapping guide. That is the paperwork an auditor typically asks to see after an incident response test.
- Detection
- T+00:00
- Declared SEV1
- T+00:12
- First external comms
- T+00:17
- Mitigation started
- T+00:50
- Resolution
- Not reached Logins were at 74.6% when the drill ended at T+01:15. Queues were still draining.
| 5. Results against the rubric (3 of 6) | Score (1–4) | Evidence |
|---|---|---|
| Detection and triage | 3 | Ana acked and checked Datadog at once, and SEV1 was declared at T+00:12, slightly after the second page (T+00:08). |
| Communication | 3 | Status updates at T+00:17, T+00:22 and T+00:34, and a proactive Enterprise email at T+00:44. The first update was vague, and the page later passed its own 03:20 promise. |
| Technical response | 2 | The web-app rollback (T+00:02 to T+00:10) and the pod restart (T+00:12) were made without verification, and the restart worsened impact (41.5% to about 5%). |
| No. | 8. Improvement actions (2 of 7) | Owner | Due | Status |
|---|---|---|---|---|
| 1 | Alert on expiry of every certificate in the chain, including root and intermediate CA, at 30 and 7 days, paging on-call rather than only Slack | Platform | 16 Oct 2026 | Open |
| 2 | Break-glass contact tree: a fourth named approver, phone numbers for all approvers and a 2 a.m. call order | Security | 13 Oct 2026 | Open |
The Game Master writes the report. The signature is yours.
Read the full sample reportSee the framework mapping
Designed to support evidence for
Note: Calmdrill is designed to support your evidence. Whether an exercise satisfies a particular requirement is always your auditor’s call. We will never tell you otherwise.
White sheet. For everyone. The scenario index.
Twelve scenarios
Written by people who have held the pager. Reliability, security and continuity, so the list goes well beyond ransomware. Each runs for 35 to 75 minutes and works in all three modes.
| Code | Scenario | Type | Difficulty | Runs for | Best for |
|---|---|---|---|---|---|
| CD-01 | The 02:14 Certificate | Reliability | Difficulty 2 of 4 | 35–50 min | First drill for any team. On-call practice. Anyone who has ever said ‘cert-manager handles that’. |
| CD-02 | Friday Flag | Reliability | Difficulty 3 of 4 | 40–55 min | Teams who ship behind feature flags, believe that makes every change reversible, and have a launch date on the wall. |
| CD-03 | Region Grey | Reliability, continuity | Difficulty 4 of 4 | 50–70 min | Teams with a DR region they’ve never really failed over to, and a BCP that says ‘RPO 1 hour’ without saying how. |
| CD-04 | The Restore That Wasn’t | Continuity | Difficulty 3 of 4 | 45–60 min | Teams who say ‘we have backups’ but whose last restore test lives in a ticket rather than a calendar. |
| CD-05 | Somebody Else’s Outage | Reliability | Difficulty 2 of 4 | 35–50 min | Any team with a synchronous third-party call in a critical path, or that has said ‘it’s their outage, nothing we can do’. |
| CD-06 | The Quiet Corruption | Reliability | Difficulty 3 of 4 | 45–60 min | Teams who own money-shaped data and whose dashboards stay green while the numbers are wrong. |
| CD-07 | Keys in the Wild | Security | Difficulty 3 of 4 | 45–60 min | AWS teams with an IAM user older than their CTO, and anyone who believes ‘AWS quarantined the key, so we’re fine’. |
| CD-08 | The Friendly Admin | Security | Difficulty 4 of 4 | 50–75 min | Security, IT and leadership together, at any company that sends invoices or is paid by bank transfer. |
| CD-09 | Poisoned Package | Security | Difficulty 4 of 4 | 50–75 min | Front-end, platform and security engineers together, especially teams who think ‘we use a lockfile, so we’re safe’. |
| CD-10 | The Open Door | Security | Difficulty 3 of 4 | 50–75 min | Exec teams, security leads and counsel who must make a breach-notification call on half the facts, best run in Leadership mode. |
| CD-11 | Nobody Home | Continuity | Difficulty 3 of 4 | 45–60 min | Teams that put everything behind SSO, and the platform and security leads who own break-glass: the credentials are in the safe, right? |
| CD-12 | The Helpful Agent | Security | Difficulty 4 of 4 | 45–60 min | Any organisation that has given an AI assistant tools: security, platform and support leads, and the exec who signed off ‘v2’. |
Green sheet. The order copy. For whoever holds the budget.
Prices
Pay once. Run as many drills as you like, across your whole organisation. The Kit costs less than an hour of consultancy.
Starter
See if it works for your team.
- The Game Master
- Scenario CD-01: The 02:14 Certificate
- After-Action Report template
- Set-up guide for Claude and ChatGPT
Kit
Everything to run a year of drills.
- All 12 scenarios: reliability, security, continuity
- Game Master for Claude, ChatGPT and a Claude Skill
- Team, Solo on-call and Leadership modes
- Evidence pack: exercise plan, attendance record, report, action tracker
- Framework mapping guide (SOC 2, ISO 27001, PCI DSS, DORA, NIS2, HIPAA)
- Facilitator Guide PDF, including a no-AI fallback
- Free updates to v1.x
$149 one-off
Whole-organisation licence
Kit, $149: on sale shortlyPro
Drills built around your systems.
- Everything in the Kit
- Scenario Forge: your AI interviews you about your architecture and risks, then writes bespoke scenarios in the Calmdrill format
- Annual Drill Programme: 12-month plan, calendar file, board summary template
- Leadership briefing cards and a comms template library
$349 one-off
Whole-organisation licence
Pro, $349: on sale shortlyPractitioner
For vCISOs, MSPs and consultancies.
- Run paid Calmdrill exercises for unlimited clients, with your own branding on the reports
- Everything in Pro
- A delivery playbook with three packages, an engagement-letter schedule and a client report cover
$990 one-off
Unlimited client engagements. The Practitioner Guide suggests $1,500 to $3,500 for one client tabletop.
Practitioner, $990: on sale shortlyNot ordering today
Try it first. The drill takes five minutes in your browser and the scorecard takes three. Neither asks for a sign-up.
Secure checkout by Payhip. Pay by card or PayPal. Your files are ready to download as soon as you pay.
White sheet. For everyone. The questions buyers ask.
Questions buyers ask
Which AI do we need?
Any paid plan of Claude (Pro, Max, Team or Enterprise) or ChatGPT (Plus, Pro, Business or Enterprise) works best, because you can use Projects to keep the files together. On free plans, use the single-file versions included in the kit. Other capable models work too.
Is the AI any good at this?
Current models are very good at playing a consistent role when they are given a fixed hidden truth, realistic signals and strict rules, and that is what each scenario provides. We playtest on Claude Sonnet, and on small models to find the floor, and publish the report from one playtest so you can judge for yourself. Use the most capable model on your plan. Small or fast models run the drill but give more away.
Will this get us through our SOC 2, ISO 27001 or PCI audit?
It is designed to help you produce the evidence auditors typically ask for when they check that your incident response or continuity plan has been tested: a plan, a record of who took part, a report and tracked actions. See audit evidence.
Note: Whether it meets a specific requirement is your auditor’s decision. We will never claim it makes you compliant.
What about our data?
The drills run inside your own AI workspace. We never see them. The scenarios are fictional, and the Game Master reminds players not to paste real secrets or customer data. If you re-skin a drill onto your own stack, a few lines of non-sensitive description is all it needs. For company use, we recommend a business or enterprise AI workspace.
Can we run it on our real architecture?
Yes. Tell the Game Master about your product, cloud and key services in a few lines and it re-skins the scenario with your service names, your tools and your numbers, keeping the same underlying failure and learning objectives. Pro adds Scenario Forge, which builds new scenarios from your own architecture and risk register.
How is this licensed?
One purchase covers everyone in your organisation, with unlimited drills. You can’t resell it or use it to deliver paid exercises to clients. For that there is the Practitioner licence, which covers unlimited client engagements with your branding.
Who makes Calmdrill?
Sandpiper Technology, a platform, cloud and operations consultancy that works with organisations where reliability isn’t optional. Calmdrill packages the way we run incident drills into something any team can use without us in the room.
Do you offer invoices, purchase orders or other currencies?
Checkout is run by Payhip, which emails your receipt and adds VAT where it applies. Prices are in US dollars, and your bank converts if needed. If you need a purchase order or a supplier form, email hello@calmdrill.com.