Lesson 027 · Phase 1, Foundations

Failure Is Normal: Designing for Partial Outages

Why most outages are partial, how to choose in advance which features may disappear, and why a degraded mode nobody has ever run is a guess.

20 min read

Lesson 27 · 27 published · 90 planned

On this page
The systems in this lessonUsed here: Marlow Books, Stagefront, and Galewatch.

Made up for this course and reused from lesson to lesson so their numbers become familiar. None of them exist. All three

Marlow Books · A small online bookshop
Four people, one server and one Postgres database. About 40 requests a second on a normal day and ten times that in the week before Christmas. The one box that the early lessons stress until it breaks.
Stagefront · An event ticketing service
Quiet most of the time, then a stadium show goes on sale at 10:00 and two hundred thousand people press the same button in the same minute. Oversold seats are a lawsuit, so correctness matters as much as speed.
Galewatch · Telemetry for wind farms
Nine hundred turbines, a reading every two seconds, over links that drop for hours in bad weather and come back with a backlog. Dashboards that lag by seconds, reports that scan a year.

Marlow Books is a four person online bookshop that exists only in this course. At five past two on the second Tuesday in September, a month after it got metrics, its payment provider stopped answering.

Not refusing. Refusing would have been a kindness: it accepted every connection the shop opened and then said nothing, which is the failure lesson 019 caught it in once.

What the shop did about that, faithfully, for the next forty five minutes: a customer pressed buy, the order row was committed as pending carrying its idempotency key, the card call went out with no transaction open, and ten seconds later the HTTP client gave up. That is the number lesson 019 found sitting there and lesson 020 argued should be two. So the customer watched a spinner for the length of a lift ride and got an error page, and the shop, with no idea whether the money had moved, left the row saying pending for the sweeper.

The sweeper is the good part of that design and it jammed. It runs every two minutes over rows pending for longer than two minutes, asks the provider what happened to each key, and completes or refunds. Ask a provider that does not answer and every question costs ten seconds. Marlow does about forty requests a second on an ordinary afternoon, and at the conversion rate lesson 009 measured at Christmas, one order every five seconds against four hundred a second, that is an order every fifty seconds. Twelve minutes in, a dozen rows were on the list, and a dozen serial questions at ten seconds each is a hundred and twenty seconds, so a job scheduled every two minutes had stopped fitting inside its own schedule.

Every book page in the shop, meanwhile, was perfect.

The founder saw it at twenty past two, on one of the graphs lesson 026 put up in August, and only because those graphs happened to be open. That is not an alert, and fifteen minutes is what not having one cost. Then they went looking for the plan, because there is one. Lesson 003 wrote it down: serve the catalogue and search from cache, disable checkout, say so honestly on the page. About 95 percent of requests served and 0 percent of the revenue.

The plan is a sentence.

Not a flag, not a button, not a switch anybody can reach. A sentence, in a runbook, written three years earlier by a person who then spent three years improving the shop.

So they did it by hand, which took twenty eight minutes: the checkout page turned into a polite refusal, a line added to the book page template, then a deploy at twelve minutes to three, which takes the site down for about ninety seconds because lesson 003 established that every Marlow deploy does. The most dangerous routine act in the shop, performed under time pressure by the only person who can perform it, as the recovery step. And nine edges went on handing out the old page anyway, because lesson 022 caches that HTML for ten minutes from the moment each city first asks, so a shop that could not take money advertised itself as open until three o'clock.

The provider answered again at twenty five to four. Ninety minutes of checkouts, which is exactly what the podcast Friday cost the shop in lesson 001, with nobody's page ever getting slower.

The unit of an outage is a feature

A partial outage is one where some of what your product does still works, which is nearly all of them. Whole systems do go dark, usually because of one shared thing underneath, but the ordinary Tuesday failure takes a slice: one dependency, one region, one table, one feature.

Lesson 024 drew the ragged version of this and handed it here. Galewatch, which collects a reading every two seconds from each of nine hundred wind turbines, cut its data across eighteen machines by farm, so when one machine goes quiet nineteen farms answer in under ten milliseconds, one farm gets nothing, and the fleet roll-up, which fans one query out to all eighteen and waits, returns nothing at all. Seventeen eighteenths healthy, and the screen everybody watches is dead.

Count that in machines and it reads as a 94 percent day. Count it in features and one of the handful this product has is gone. Nobody buys machines.

So the unit of an outage is the feature, and the number that matters is not how much of your fleet is up but which of the things you do still happen. There is a design lever inside that, upstream of anything lesson 026 taught about displaying the result. A fan-out has to decide, in advance and in code, whether a missing leg is an error or a footnote. Wait for all eighteen and one quiet machine deletes the view for everybody. Answer with what came back, labelled nineteen of twenty, and one farm's engineers lose their farm while everybody else keeps working. Both are defensible. Only one is a decision, and the default is the other one.

Which is the shape of the whole subject. Every system already has a degraded mode. Marlow's on that Tuesday was "the checkout takes ten seconds and then apologises", chosen by the value of one HTTP client's timeout, a number lesson 020 established nobody chose. If you do not choose your degraded mode, your timeouts choose it for you. They choose badly, because a timeout knows nothing about what your product is for.

Three modes, not one per dependency

Writing the list looks like more work than it is. Start from the dependency rather than the feature. Marlow has more dependencies than this and five that decide its modes: the payment provider, box A's Postgres, the replica it bought in June, lesson 008's Redis and lesson 022's content delivery network.

If this goes dark What stops Which mode
Payment provider taking money no money
Box A's Postgres every write, sign in, checkout no writes
The replica sign in, baskets, search, prices none, if you fall back
Redis or the CDN nothing, until the ceiling no cache

In words, because that table is the plan and the spoken version skips tables. The provider going dark stops the shop taking money and nothing else, which is September. Box A going dark stops every write, so nobody signs in, checks out or posts a review, while book pages and search carry on from the copy. The replica going dark stops sign in, baskets, search, the publisher page and the live price on every book page, which sounds like the worst of the five and is the one you can delete. The caches going dark stop nothing until the request rate is above what the database can serve alone, at which point they stop everything.

Five dependencies, three modes, because Redis and the CDN fail into the same one and the replica need not have a mode at all. Modes collapse, which is the useful discovery in this exercise and the reason it terminates: a shop with fifteen dependencies has about four degraded modes, which is short enough to put on one page and argue about with somebody who does not write code, who should also settle the order, because it is a revenue question and not an engineering one.

Now the deletion.

Lesson 011's routing audit sends four kinds of read to the replica: full text search, the uncached publisher page, the book page's stock and price, and session lookup by id. That last one is why a dead replica signs everybody out, and lesson 007 is why it matters twice, because the session row holds the basket, so no session means no basket means no checkout while the machine that takes the money sits there healthy. Four features and all of the revenue, gone with a copy lesson 011 priced at 6.5 percent of one box and said was never bought for throughput.

Every one of those reads exists on box A too, because the primary is where they came from, so the fallback is two lines: try the replica, and on failure ask the primary.

Lesson 011 already wrote down the direction, for a different reason: it made the primary the default and the replica the thing a developer has to type, because somebody who forgets gets a query that is slower, correct and boring, which is the right direction to fail in. Widen that from a developer forgetting to a machine being dead and you have a rule. Fallbacks run one way, toward the copy that decides. Falling back from a replica to a primary costs capacity. Falling back the other way, from the primary to a copy, costs correctness, and lesson 009 already paid that bill: a cached stock number sold forty copies of a book the shop had twelve of, and the packing room found out by reaching for the thirteenth.

The bill for the safe direction is real and must be measured. Marlow's is 1.04 seconds of database work a second at the Christmas peak, lesson 010's figure and lesson 011's six and a half percent of sixteen cores, so box A can absorb every read the replica was doing. What it cannot absorb is the connections, since lesson 010 already reserves 96 of Postgres's 100 and lesson 011 added a second pool per process on top. At a shop whose replicas carry real load this is no fallback at all, it is a stampede pointed at the one machine you have left.

The ceiling you fall through

The mode nobody plans for is the one where nothing is broken.

Take Redis away from Marlow during Christmas week and nothing errors. Every book page still renders from Postgres, at the five milliseconds lesson 001 measured as the cost of a page before the cache existed. Lesson 008 did the arithmetic: the shop is offered 384 page requests a second against a pair of boxes that can serve about 260 uncached, which is 148 percent of the ceiling. Losing the CDN is the question lesson 022 sent here by name when it priced a CDN at 53 minutes a year, and it lands softer, because Redis is still standing behind it and the origin only goes back to the load it carried before March.

What happens at 148 percent is not a slightly slow shop. Lesson 001 measured the shape at a worse ratio: the podcast Friday offered just over 400 requests a second to a box serving about 200, a book page took eleven seconds by 10:47, and at 10:52 a health check restarted the process and killed the checkouts in flight. Everybody waits and nobody buys.

At Christmas the shop serves 400 requests a second and packs an order every five seconds, so one request in two thousand is an order. All of the revenue is a rounding error in the request count. Protecting it costs almost nothing, and a queue serving whoever arrived next spends 99.95 percent of a scarce resource on browsing.

Lesson 021 gave this its vocabulary, so take it as placed: load shedding is a reaction to a measurement of yourself, refusing what you can because serving nobody slowly is worse than serving most people quickly. It also named the signal, requests in flight rather than CPU, and lesson 026's USE half is where Marlow finally got an instrument that reports it.

So the design of "no cache" mode, on machinery the shop owns. Refuse full text search first, lesson 010's sixty milliseconds and 0.96 of the database's busy second. Serve book pages until requests in flight say stop, then hand the rest a static apology the CDN can serve without asking the origin. Never refuse a checkout. Lesson 021 put per-path rules into that balancer in February, so what is missing is one rule whose job is to protect a path rather than limit one, the half of lesson 021's split that had nothing to motivate it until today.

Shedding is how a total outage becomes a partial one. It is the only mechanism here that takes an outage nobody chose and turns it into a mode somebody did.

Fail open, fail closed, and the third one

When a check cannot be made, you either let the work through or you stop it. Failing open lets it through. Failing closed stops it. Every dependency you have carries that decision whether anybody made it or not, and the default is whatever your exception handler does.

Lessons 023 and 024 left the test here, and it is the same sentence both times: a read that decides is the first half of a write. Fail open on a suggestion. Fail closed on a decision.

Marlow's book page carries a stock number, which lesson 008 established is a suggestion, because clicking buy does not buy anything: the conditional update inside the checkout decides, since lesson 015. If Redis is gone, render the page with no badge and let the customer press the button. Failing closed there, refusing a page because you cannot show a badge, throws away a sale to protect nothing.

The checkout's own guard is the other kind. There is no safe default for "is there a copy of this book", and lesson 009 published what the unsafe one costs.

Then the case every textbook uses and nobody prices. Lesson 003's second check question put a fraud check at 99.99 percent onto the purchase path of Stagefront, the ticketing service in this course where a stadium show goes on sale at exactly 10:00 and two hundred thousand people press the same button in the same minute. At 99.99 percent that service is unavailable for 53 minutes a year, lesson 003's own table. Fail closed and you refuse every purchase for those 53 minutes. Fail open and you take 53 minutes of unchecked purchases.

Lesson 009's model is how you choose: the chance of being wrong, times how long you stay wrong, times what a minute costs, of which only the middle term is yours. Run it and failing open usually wins, because a fraction of the purchases in 53 minutes is cheaper than all of them.

Two things spoil that, and the first is the one people miss. Those 53 minutes are not spread evenly. A fraud service falls over when it is busy, it is busy when you are, so its failures arrive correlated with your on-sale, and lesson 003 told you Stagefront's entire business happens in 400 minutes a year. Pricing a fail-closed decision at the average minute is how you come to refuse a stadium.

The second is why fail closed keeps winning the meeting anyway. A fraud loss arrives with a receipt: an amount, a date, a name on it. A refused purchase arrives as nothing whatsoever, which is lesson 026's rule that a missing measurement and a measurement of zero look identical. The failure with a receipt beats the failure with no receipt, whatever the arithmetic says. I have lost that argument while holding the better numbers.

There is a third option and it is usually the right one: fail open and write down that you did. Let the purchase through, put a row in a table saying it went unchecked, review the rows when the service is back. That is lesson 019's pending row in a new costume, with lesson 018's warning attached, because a queue of skipped checks with no owner is failing open with extra steps and better branding.

Stagefront is where refusing is right, and it has already done it. Lesson 020's Tuesday in March: the payment provider slowed to about 900 milliseconds, the checkout tier's pool filled, a human halted the on-sale at 10:04, and it reopened at 10:40 with one number changed, after which the remaining twenty three thousand seats went in forty six seconds. The strongest degraded mode in this course is a button that stops taking money on purpose, and it works because the forty thousand seats lesson 017 took as the stadium's size do not spoil in half an hour: that show is the only place to buy that show. Ninety minutes of a closed bookshop is ninety minutes of people buying the same book somewhere else. Whether refusing is correct is a fact about your market, not about your architecture.

A switch you cannot reach

A degraded mode asks three things of its switch. Marlow's had the first by luck and failed the other two.

It must not depend on the thing that is broken. On 28 February 2017 an engineer at Amazon running a playbook mistyped a command, removed far more S3 capacity in us-east-1 than intended, and took the index and placement subsystems down for hours. The detail to carry is smaller: the AWS status dashboard could not be updated to say so, because the console that updates it depended on S3 in us-east-1. The part of a system whose only job is to work when the rest does not was sitting on the rest.

It must take effect faster than your caches, and Marlow's did not: nine edges held the old page for ten minutes after the banner went live, and the purge that clears them all is the button lesson 022 says switches your CDN off.

And it must be flippable by whoever is awake. Lesson 003 priced the 95 minutes between an alert that could have reached a phone at 10:44 and the customer email that actually arrived at 12:21, and at a four person shop the people who find out first are in the packing room with a tape gun. A switch that needs a deploy is not a switch, because nobody holding a tape gun is going to ship one.

What Marlow built that week costs nothing, because it reuses something the shop had. Lesson 022 split the book page so nothing per-person sits in the cached document, and lesson 023 pushed the price into that same small second request the browser makes on every page. It is live, uncached and already there. Two more fields in it: whether checkout is open, and a sentence to show when it is not. The values come from a file on each box rather than from a database, so they survive the mode where the database is the problem, and changing them is thirty seconds over ssh with no deploy. Two files can disagree, which is fine for a banner and a disabled button and would not be fine for anything that decides: a switch is allowed to be inconsistent in a way a stock count is not.

Last, what it says. Lesson 003 has a Galewatch engineer reading a green tile for a turbine feathered since eleven that morning, and lesson 026 widened that for dashboards: an answer that does not say what it is missing is a lie with a chart on it. A customer's screen is the same rule at a bigger radius, and "something went wrong" is the green tile of retail. A degraded mode has to name the thing it cannot do, in the words of the thing the customer came for, and say what it can still do.

The mode that rotted

Read lesson 003's sentence against the shop that exists in September, feature by feature, and it does not survive.

"Search from cache" was never true. Search has always been a Postgres full text query over 1.2 million descriptions and their reviews, it is the thing that pushed the working set past 40 gigabytes in lesson 005, and it has lived on the replica since June. Nothing has ever cached it.

"Catalogue from cache" was true when it was written and stopped being true in May. Lesson 009's split turned a book page into a cached shell plus one primary key read for the live price and stock, and lesson 023 moved the price one radius further out again. The cached document renders a book with no price.

And yet the shop can very nearly do it anyway, because lesson 011 sent the book page's price and stock read to the replica in June for reasons unrelated to any of this. If box A dies, pages and search keep working from the copy and the checkout correctly refuses, since the machine that decides is the one that died. The mode exists. It exists by accident, it is nowhere in the runbook, and a routing change made for a performance reason on any Tuesday would delete it with nobody noticing.

Nobody made a mistake here, which is worth being blunt about. Every change between the runbook and September made the shop faster or safer: a shared cache, a split page, a replica, an edge. Every improvement to the fast path is an edit to the slow path, and nobody diffs the slow path. It is the only code in your system whose bugs stay invisible until the worst day of the year, because it is the only code that never runs.

So run it. Lesson 025 put a failover rehearsal in Marlow's calendar for the first Saturday of every month, once the founder measured the rebuild at forty one minutes instead of the night the runbook claimed. A second line on that Saturday costs nothing: at forty requests a second, flip the flag for two minutes and look at your own shop. Lesson 026 coined the small version, fire every alert once on purpose, because a signal that has never moved is untested. A mode nobody has entered is worse, because somebody wrote a sentence about it and everybody believes the sentence.

The first one is in October, and at the end of this lesson it has not happened.

Then the honest limit. You can only rehearse the modes you thought of, and lesson 026 already said the next failure is in the part you have not thought about. A list of modes is not a prediction. Its whole job is to make sure nobody has to decide what the shop is for at half past two on a Tuesday, with a spinner on the screen and a customer waiting.

Recap

The unit of an outage is a feature. Galewatch losing one machine of eighteen is a 94 percent day counted in hardware and a dead product counted in screens, and nobody buys hardware. Count features, weight them by money, and decide in code, in advance, whether a missing leg of a fan-out is an error or a footnote.

If you do not choose your degraded mode, your timeouts choose it for you. Marlow's September mode was ten seconds of spinner and then an apology, picked by a number lesson 020 established nobody chose. Every system already has one; the question is whether an exception handler picked it.

Modes collapse, so the list is short. Start from each dependency, write down what stops, and watch several of them land in the same mode. Marlow has five things that can go dark and three modes, and one of the five can be deleted, because fallbacks run one way, toward the copy that decides. Falling back to the primary costs capacity you must have measured; falling back the other way costs correctness.

One request in two thousand is an order. The path carrying all the revenue is a rounding error in the request count, so protecting it is nearly free while first come first served spends everything on browsing. Shedding is how a total outage becomes a partial one: at 148 percent of your ceiling, refusing a third of the browsing leaves a shop and refusing nobody gives everybody lesson 001's eleven second page.

Fail open on a suggestion, fail closed on a decision. A stock badge is a suggestion and the conditional update is the decision. Never price a fail-closed check at the average minute of the year, because a shared dependency breaks when you are busy and Stagefront's year is 400 minutes long. And the failure with a receipt beats the failure with no receipt in every meeting, which is why failing open and writing down what you skipped is usually right, and why that table needs an owner.

A switch that needs a deploy is not a switch. It cannot depend on what is broken, which is what S3's status page did in 2017; it has to beat your caches, since nine edges held Marlow's old page for ten minutes; and the person who notices has to be able to flip it. Then it has to name what it cannot do, because "something went wrong" is the customer-facing version of a green tile on a dead turbine.

Every improvement to the fast path is an edit to the slow path, and nobody diffs the slow path. Marlow's written degraded mode aged out under three good decisions and the one it really has works by accident. A mode nobody has entered is a guess with a runbook entry, so put it in the calendar beside the failover rehearsal and enter it on purpose at forty requests a second.

Check your understanding

  1. Marlow's replica goes dark at 11:00 on the Tuesday of Christmas week. Using lesson 011's routing audit and lesson 010's connection arithmetic, say which features stop and whether falling all four reads back to box A works at that moment.

  2. Write the entry condition for Marlow's "no cache" mode as something a machine can evaluate, using lesson 008's ceiling and lesson 021's advice on which signal to shed against. Say which path it refuses first and which it must never refuse.

  3. Take lesson 003's proposed fraud check at 99.99 percent on Stagefront's purchase path. Decide fail open or fail closed, say what you would measure for a month first, and name the number that would flip your answer.

  4. Marlow's new flag lives in a file on each of two boxes, which can disagree. Name something it must never control, and which earlier lesson explains why.

  5. Take a service you have worked on and write its degraded modes, no more than four, each with the dependency that triggers it, what the customer sees, and who may turn it on. Then say which you have ever run on purpose.

Next lesson

028 API Design: Contracts That Survive Change. Every mode today had to tell a caller what it could not do, and next lesson is about the contract that message lives inside: what an interface can promise, what you may still change once somebody depends on it, and how a version you thought was internal becomes one you cannot break.

Finished reading?

Marking a lesson done keeps your place on the course index. It is stored only in this browser.

Tip: use the ← and → keys to move between lessons.