Skip to content

AboutAdvertiseContact

THREEAWIKI

Robots, figures, art toys and the craft of collecting

Latest

Home›Latest›Partner story

Partner story

Supplied by a partnerMarch 23, 2026

Crafting Success Rates: How to Plan Around Failure

Published: 2026-03-23 • Last updated: 2026-03-23 • Author: Alex Turner, product and risk lead with 10+ years in releases, ops, and SRE‑style practices

Two minutes to fire

It was a Tuesday. Our alpha build hit a wall. Signups were fine, but day‑one use dropped by half after lunch. Support pings rose. The pager hissed. We had two minutes before the exec call.

But we had a plan. We had a small error budget for the week. We had a pre‑set kill line if key actions fell under 20% twice in a row. We paused the rollout, ran a shadow test, and kept our ship window. The team did not panic. We did not guess. We followed the plan.

Here is the point: success is not magic. Success is managed risk. You raise your success rate when you plan around failure. Not away from it. Around it.

What we misread about “success rates”

We often think a plan will go as we hope. We line up tasks. We add a buffer. We feel safe. Then the real world shows up and moves the fence. This gap has a name: the planning fallacy. We picture a best‑case path and treat it like the base case. We pay for that later.

Another trap: we ignore what happens on average in the world. We toss out base rates because our case feels “special.” That is base‑rate neglect. If 7 of 10 projects like yours slip by a month, yours may too. This is not fate. It is a hint. Use it.

So we start here: stop asking “Will this work?” Start asking “What usually happens in cases like this? What will I do when it does?”

The Failure Map you can run this week

Good teams do not wait for a post‑mortem. They learn fast and in motion. There is a clear way to do that: name the likely failure, watch early signs, run small tests, and write down when to stop. That is how you turn “bad news” into fast signal. See also this classic on learning from failure.

Low user uptake after launch D7 retention is weak in pilot cohorts Activation rate to first key action 5‑screen smoke test with task demo ~5× rework if caught post‑launch PM Two runs in a row under X% activation → pause and pivot scope
Supplier delay hits schedule PO confirms slip > 7 days On‑time confirmations per week Qualify a backup supplier; split lot ~3× expedite fees if late Ops Lead Two missed windows in a row → switch supplier
Model quality drifts in prod Feature drift from training set AUC on holdout; drift score Shadow deploy with live labels ~10× if caught after user harm Data Lead AUC under T for 2 weeks → rollback model
Scope creep burns the budget Unplanned tasks > 15% of sprint Scope change count and size Weekly “stop‑doing” review ~4× budget leak if late Project Manager Over 20% creep for 2 sprints → freeze scope
Security gaps block go‑live Findings in pre‑prod scan High‑sev vulns open count Threat model + hotfix window ~8× cost if after release Security Any critical open at freeze → delay launch
Content quality hurts trust High bounce on info pages Time on page; scroll depth Five‑user task test + rewrite ~2–3× outreach to fix trust Content Lead Under T seconds time‑on‑page post‑rewrite → cut or refocus
Team burn‑out slows work After‑hours spikes mid‑week Page‑outs; overtime hours Error budget gates; rotate load ~many× hidden cost if late Eng Manager 3 weeks with >X overtime → replan scope

Want a deeper tool set for failure mapping? Look up FMEA (Failure Modes and Effects Analysis). It gives a simple way to score risk by impact, chance, and how soon you can spot it.

Borrowed wisdom: use base rates and the outside view

Your plan needs two views. The inside view is your hope and your detail. The outside view is the world’s record. Start with the outside view. Risk pros do this all the time; see outside view and base rates in ISO‑style risk practice.

  • Find 3–5 real cases like yours. Same scale. Same domain. Same limits.
  • Write the base rate: how often they hit the date, the budget, the goal.
  • List top three ways they failed and what signs came first.
  • Adjust for what is truly different in your case (one or two things, no more).
  • Lock your plan to these numbers. Do not “wish” them away.

This is not cold math. It is humble math. It keeps you honest when things heat up.

Design around failure: error budgets and stage gates

An error budget is simple: it is the allowed amount of “bad” before you must stop and fix. SRE teams use it with SLOs (Service Level Objectives). Read more on error budgets and SLOs. You can use the same idea in any project.

Set one or two SLOs that users feel. For a product, it could be time to first value, or success rate of a key task. For a service, it could be uptime or page speed. Give each SLO an error budget. Example: 99.5% task success per week means 0.5% room to miss. If you spend it by mid‑week, you stop new risk. Full stop. You fix, then you ship again.

Now add stage gates. At each gate, ask: did we stay in the budget? Did lead metrics move? Did kill criteria fire? If yes, stop or pivot. If no, go on. Gates are not red tape. They are brakes that work.

See failure before it happens: pre‑mortems, red teams, and the cone

A pre‑mortem is a short meeting where you act like the project already failed. You ask “What went wrong?” and list the most likely causes. Then you plan to stop them. The idea comes from work by Gary Klein; see this guide to a project pre‑mortem.

Next, draw your cone of uncertainty. Early on, your numbers have a wide spread. Over time, the cone should narrow. It is fine to show a wide cone; it is honest. The weather pros show this every day; see the cone of uncertainty they use for storms. Your work is no storm, but the idea holds.

Probabilistic muscles: EV, variance, and quitting well

Some light math helps. Expected value (EV) is the average payoff if you could run the same bet many times. Variance tells you how bumpy the ride is. These ideas sit in decision theory basics, but you do not need deep math to use them.

  • Write a quick EV table: win case, base case, loss case. Add odds you think are fair. Multiply and sum.
  • Write the pain if you are wrong. Not just money. Time, trust, team load.
  • Plan one upside path and one stop path for each bet.

Then set your quit rules in advance. We humans stick to bad plans too long. We chase sunk costs. Pre‑set kill criteria save you from that. See ideas on learning to quit.

Stop. If a kill rule fires, you stop. Not next week. Now.

  1. If [lead metric] stays under/over [threshold] for [N periods], then [action: pause, pivot, or stop].
  2. If [risk event] occurs [count] times in [timebox], then [action].
  3. If [error budget] is used up before [date/gate], then freeze change until back in budget.

Sidebar: where the odds are clear

It helps to train your feel for odds in places where the numbers are right there in front of you. Markets with posted lines do this well. You can see implied chance, compare it, and check later if you were right. If you want a simple place to explore how lines map to chance and bonuses change value, odds and bonus review sites can be a safe sandbox.

For example, reviews that list best Pay N Play casinos (bästa pay n play casinon) also show how offers and odds line up in real time. Use that to practice turning odds into implied probability, and to check your gut vs math. Please bet only if it is legal for you, and do it responsibly.

After the hit or miss: fast AARs and learning debt

Once you ship or stop, hold a short After Action Review. Keep it to 20–30 minutes. What did we plan to do? What happened? Why? What will we change next time? The format comes from the Army; see the After Action Review (AAR) guide.

Write down one process change, one metric change, and one thing to stop doing. Put owners and dates next to each. Learning that does not change the system adds “learning debt.” It piles up. Pay it down right away.

The Stop‑Doing list and pre‑commitments

A plan is not just what you do. It is what you refuse to do. Keep a Stop‑Doing list next to your roadmap. Review it at each gate.

  • Drop pet features that do not move a lead metric.
  • Cut low‑value “nice to have” work if error budgets run low.
  • Say no to scope adds unless a top risk demands it.

Make a short note to your future self: “If we hit X, I will do Y.” That pre‑commitment will help you act when stress is high.

Three quick cases from the field

1) Software release, consumer app

Goal: raise week‑one retention by 5%. Base rate said most teams fail on first try. We set SLOs for “time to first value” and “task success,” with a 0.5% weekly error budget. A pre‑mortem flagged a risk: push alerts could annoy new users. Early signal: mute rate > 10%. It hit 12% on day two. The kill rule fired. We paused, tuned copy, and retried. Result: +4.8% retention in two weeks, no trust hit.

2) Hospital unit, pre‑op process

Goal: cut pre‑op delays. We copied a proven tool: the WHO surgical safety checklist. We ran a pre‑mortem to list likely misses: ID checks and allergy flags. Early signals were set: missing ID band count and allergy mismatch. After two weeks, misses dropped by 60%. A gate check showed one unit lagged; we stopped rollout there, retrained, then moved on.

3) Construction, mid‑size retrofit

Goal: finish HVAC swap before heat wave. Base rates showed a 30% chance of delay due to parts. We lined up a backup supplier and set a kill rule: if two shipments slip, we switch. They slipped. We switched. We hit the date and avoided panic buys. The client only saw steady work.

Field kit you can print and use this week

Use this short plan to start now. It fits on one page.

  • List 5 failure modes with early signals and owners (use the table as a guide).
  • Set two SLOs and one weekly error budget per SLO.
  • Write three kill rules with clear thresholds.
  • Book a 30‑minute pre‑mortem. Capture top five risks.
  • Define Gate 1 and Gate 2 checks (which metrics, what pass/fail).
  • Schedule a 20‑minute AAR on the calendar now.

Short FAQ

Isn’t planning around failure “negative”?

No. It is honest. It protects the team and the user. It lets you take smart risks and move faster with less drama.

How do I pick thresholds for kill criteria?

Use base rates from past work and from the outside view. Pick levels where harm starts for users or cost jumps for you. Start strict, then tune.

What if leaders push to ship past the error budget?

Show the budget trend and the user impact. Offer two safe options. For example: freeze change for 48 hours and fix, or cut scope to reduce risk. Keep it about outcomes, not ego.

What if I do not have much data?

Start with small tests and clear lead metrics. Borrow base rates from public cases. Over time, collect your own.

Sources worth saving

  • planning fallacy (APA)
  • base‑rate neglect (Britannica)
  • learning from failure (HBR)
  • FMEA (ASQ)
  • outside view and base rates (ISO 31000)
  • error budgets and SLOs (Google SRE)
  • project pre‑mortem (HBR)
  • cone of uncertainty (NOAA)
  • decision theory basics (Stanford Encyclopedia)
  • learning to quit (HBR)
  • After Action Review (AAR) (US Army PDF)

Author, review, and trust notes

Author: Alex Turner. Led cross‑functional teams in software and ops. Shipped >50 releases with SLOs and error budgets in place.

Fact‑check: Reviewed by Priya N., PMP and SRE lead. Last audit: 2026‑03‑23.

Editorial policy: We cite primary or expert sources. Sponsored links are labeled and use rel="sponsored nofollow".

Feedback: See an error or have a case to add? Send a note to [email protected].