Lesson 038 · Phase 2, Mechanisms

Event-Driven Architecture and Its Failure Modes

A producer that stops knowing who is listening also stops knowing what a replay will do, and every new consumer brings its own duplicate problem.

22 min read

Lesson 38 · 38 published · 90 planned

On this page
The systems in this lessonUsed here: Marlow Books, Stagefront, and Galewatch.

Made up for this course and reused from lesson to lesson so their numbers become familiar. None of them exist. All three

Marlow Books · A small online bookshop
Four people, one server and one Postgres database. About 40 requests a second on a normal day and ten times that in the week before Christmas. The one box that the early lessons stress until it breaks.
Stagefront · An event ticketing service
Quiet most of the time, then a stadium show goes on sale at 10:00 and two hundred thousand people press the same button in the same minute. Oversold seats are a lawsuit, so correctness matters as much as speed.
Galewatch · Telemetry for wind farms
Nine hundred turbines, a reading every two seconds, over links that drop for hours in bad weather and come back with a backlog. Dashboards that lag by seconds, reports that scan a year.

Marlow Books is an online bookshop run by four people, and it exists only in this course. On the third Tuesday in April its orders graph said a hundred and nine orders had been placed in one minute. One had.

The graph was a fortnight old. Lesson 026 argued for it in August, as an orders a minute line with the same hour last week drawn behind it, and priced it at about four minutes of work, because the shop has counted orders in a table since the day it opened and a graph off that table is a count and a group by. Then the founder did not build it for eight months. Lesson 026 blamed that on the number belonging to nobody rather than on anything being hard, and what finally got it drawn was a new way of doing it badly.

Since lesson 018 the broker has carried exactly one kind of message, written by the checkout when an order is paid for and read by one consumer that sends the confirmation email. On the first Tuesday in April the founder renamed that message from send_confirmation to order_placed and hung two more consumers off it.

Twelve lines each. That part is worth noticing before the failure.

reorder_watch reads the stock of the title that just sold and emails the two buyers from lesson 010 when the shelf is down to under three copies. The buyers had asked for that on the wet Saturday in January that lesson 030 opens with, with a finger on a diagram. order_count adds one to a counter. The checkout did not change, and nothing in it knows either consumer exists, which was the whole appeal.

Two weeks later, on the third Tuesday, the founder built something lesson 018 had been owed since Christmas: an alert on the dead letter queue being non-empty at all.

It fired the moment it went in. The drawer held a hundred and eight messages and had held them since the Thursday evening in March that lesson 036 is about, when a slow email provider got every one of those orders emailed three times and then filed all hundred and eight under failure. The founder knew they were done; they had spent a Friday morning proving it.

The only way to make an alert on a non-empty queue go quiet is to empty the queue. Lesson 018 left a five line script for putting dead messages back on the main queue, written in Christmas week for forty messages that genuinely had not been processed. At 11:46 the founder ran it.

Lesson 036's whole argument is that the drawer holds the messages the broker gave up on, which is not the same set as the work that did not happen. The founder knew that. The script didn't.

A hundred and eight messages went back on the queue, and all three consumers were finished with them inside the same minute. Lesson 036 measured the email handler at a 240 millisecond median and gave it four workers, so a hundred and eight is twenty six seconds of work shared four ways: six and a half seconds, with every one of them calling the provider.

March's bug was fixed in March, the week after the Thursday. The email consumer now claims an order with a row carrying a null sent_at before it calls the provider, stamps that row once the provider answers, and a second delivery that finds a stamped row acks and sends nothing. It works. It has worked on every order since.

It could not work on these. The hundred and eight in the drawer are the orders the fix exists because of, which puts them on the wrong side of it. They went out before there was a claim row to write, and they are the only messages in the shop that are both older than the fix and still replayable. The consumer found nothing, wrote a fresh claim, and sent.

A hundred and eight people got a confirmation email for a book they had bought a month earlier.

Everything in a dead letter queue is older than your most recent fix. That is the one thing you can say with confidence about every message in one, and it is worth saying before you empty one. The March fix answers one question: is another worker sending this message right now. Nobody designing it was thinking about April.

reorder_watch evaluated all hundred and eight against April's shelf and emailed the buyers about every title under three copies today. Nobody recorded how many that was. By the afternoon one buyer had filtered the lot into a folder they do not read, which is the real cost and appears on no graph.

order_count added a hundred and eight to the 11:46 bucket.

The founder saw the spike first, inside the minute, and had lesson 001's reflex, which is fair for a shop that was once named on a podcast at 10:40 on the Friday before Christmas. They opened the orders table. One order, 11:46, one book. The replies from confused customers started arriving about an hour later and took the rest of the day.

So which number was wrong? The one on the graph the shop had just started to trust, and it had been counting the wrong thing from the minute it was written: not orders, but deliveries of a message about orders, and lesson 036 spent a whole lesson proving those are two different numbers.

A producer that publishes a fact stops knowing who is listening

Event-driven here means a service announcing that something has happened and not knowing or caring who acts on it. The alternative, which is what nearly every service you have written does, is calling another service and waiting.

An event is a statement that something happened. Past tense, no addressee, true whether anybody reads it or not. A command is an instruction to one named recipient to do something that has not happened yet. order_placed is an event. send_confirmation is a command, and it sat on a queue for four months while the shop thought of it as the place where orders go. Nothing in the plumbing tells you which one you are holding: same broker, same bytes, same acknowledgement lesson 036 defined. The difference lives entirely in who is allowed to care.

The trade is one sentence and the rest of today is its consequences. A producer that publishes a fact stops knowing who its consumers are.

That is the selling point. The founder added two consumers in twelve lines each and never opened the checkout. Had the checkout called all three instead, each would have joined the request path, and lesson 003's arithmetic is the one to run on it. That chain is Stagefront's ticketing checkout and lesson 017 took the trouble to say it is not Marlow's: five components between 99.99 and 99.9 percent multiply out to 99.78 percent and 19.3 hours a year. Borrow the shape and not the number. What transfers is the direction, and the direction only goes one way. Lesson 017 refused to put the card call on a queue because the customer is standing there and the answer is the product. A reorder notice is the opposite case. Nobody is standing there, one that arrives four seconds late is indistinguishable from one on time, and one that fails should never cost a sale.

It is also the source of every failure below. What you give up divides into six things. Who is listening, which the broker knows and your diagram does not. What the event means, which the payload decides and the name does not. When it happened, against when a consumer reads the database. How many times, which lesson 036 settled for one consumer and nobody re-settles for three. What shape it is, which is lesson 052's subject. And where the call stack went, which is the hard one.

Two problems today states and leaves. Marlow's checkout commits the paid transaction and then publishes the message, which is two machines and therefore a gap: the order is paid for, no message exists, nothing notices, and the customer never gets an email. That is the write to two places problem, and lesson 040 owns it. Once three consumers each retry against a dependency of their own you want fuses, which is lesson 042.

The event that is a command in disguise

Marlow's message carried what the email template needed, and renaming it did not change a byte.

order_placed, as the broker had carried it since December
  to      the customer's email address
  title   the book, as the email prints it
  total   "14.00", formatted for the email

Three fields, every one of them there for the email. reorder_watch needs a stock count, and stock is keyed by ISBN, which is how lesson 008 keys the page cache too. So the first Tuesday's twelve lines were really thirteen: the founder had to add the ISBN before a second consumer could do anything useful with the message.

That is the diagnosis, and it is cheap to run on your own system. If adding a consumer forces you to change the producer, what you had was a command. An event that only makes sense to one reader is addressed to that reader whatever the type field says.

Growing a payload has a published price at this shop. On the Friday of Christmas week at ten to five, lesson 018 has the founder adding a field to the confirmation payload in producer and consumer at once, with forty old-shape messages still on the queue; the new consumer raised a KeyError on each, and eleven minutes went by before lesson 017's age alert fired. April cost nothing, because at lesson 027's one order every fifty seconds, a figure nobody at the shop has ever checked, four workers leave the queue empty at 11:20 on a Tuesday. The shop got away with the exact change that cost it eleven minutes in December, and the reason is the calendar rather than anything it learned. Lesson 052 is where that stops being luck.

The over-correction is the first thing everybody reaches for. Put the stock count in the event and reorder_watch never touches the database. Now the checkout is responsible for knowing what a consumer needs, which is the coupling you just paid to remove, and the number is stale the instant it is written. An event's payload should be what you would want in a log line a year from now about that fact. Lesson 059 is where that stops being a style preference and becomes a record somebody has to keep.

An event is a fact about the past and your database is the present

Lesson 035 moved Marlow's stock decrement into the same commit as the pending order row and its idempotency key, so a copy leaves the shelf at the moment of intent. The card is charged afterwards with no transaction open, and a second short transaction moves the row to paid. The message goes out when that commits.

The shelf moves first. The event exists second. Between them sits a payment provider, which lesson 015 put at about 300 milliseconds on a normal evening. Lesson 019's bad case is longer than people expect. A row has to sit pending for longer than two minutes before the sweeper will look at it, and the sweeper runs every two minutes, so close to four minutes can pass before the job even reaches the row. Then the job asks the provider what happened to that key, and lesson 027 watched a dozen of those questions at ten seconds each come to a hundred and twenty seconds, which is the sweeper's entire schedule. The event for an order whose card call hung arrives on the far side of all of that.

reorder_watch reads the shelf when the event arrives, so it is always reading a row that is ahead of its own event.

That is the opposite of the failure a reader is braced for. Lesson 012's replication lag has the consumer reading a copy that is behind, and there is no replica anywhere in this one. Point that read at the June replica, where lesson 011 sends the book page's stock read, and you get both errors at once in opposite directions, with lesson 012's measured 1.1 second daily peak between them.

Lesson 011's audit would send this read to the replica without pausing, because the question it asks of a read is whether anything decides on it, and an email to a buyer decides nothing. That audit has a column for what a read is allowed to be wrong about and no column for which moment it is asking about. Marlow's consumer reads box A, for no reason anybody wrote down.

Now the part that costs more than the lag. Lesson 035 also built a refund path, whose job is putting a held copy back on the shelf when a card call fails. It publishes nothing. So the message stream carries orders that were paid for and says nothing at all about copies that came back, and anybody deriving the shelf from those messages will be confidently wrong in one direction forever.

If the events do not cover every change to a thing, you cannot derive that thing from the events. Written down it sounds too obvious to say. It is the most expensive mistake in this architecture, because the first few consumers always can derive it, right up to the day somebody adds the path that publishes nothing.

Ordering arrives here as lesson 037's bound with the guarantee taken out. Inside one of 037's partitions the order of records is real. Marlow's queue has no partitions and four workers on two boxes, so two events about the same title can be handled in either order, for lesson 034's reason: nothing talks, so nothing is ordered. At one order every fifty seconds that will happen rarely enough that no test finds it and a customer does. The only real fix is 037's, which is to partition by the key whose order you need and accept that the busiest key becomes your ceiling.

The architecture nobody decided to build

Galewatch is a company in this course and nowhere else: it takes a reading every two seconds from each of the wind turbines on the farms it monitors, and sells the owners dashboards. Four ingest boxes append every reading to the log lesson 037 built. Six writer processes read that log and put each reading into Postgres. Since the Monday after 037's Thursday, a second pair of processes reads the same records under a group id of its own, looking for a turbine sitting at zero while its neighbours turn.

Nobody at Galewatch has ever said the words event-driven architecture. They have one, and it arrived as a configuration change.

Amplification shows up here with real numbers rather than Marlow's three. Lesson 036 derived 620 records a second from 1,240 turbines at a reading every two seconds. Two groups make that 1,240 handler runs a second off 620 appends, and a third group would be another 620 that the ingest boxes would never see. The log's own cost does not move, which is 037's point about a reader being a number. What moves is the count of places that can be slow, wrong or behind, and nothing warns you when it goes up.

The failure that belongs to Galewatch is the one no single service can see. On the Thursday the fix was to move the writer group's position back to 11:00 and let it re-read twenty one minutes, and 037 is explicit about what that did and did not do. A replay moves one group's offsets, so the readings came back and the alarms never did. A group with a new id starts at the end of the log by default, so the watcher has still never looked at those fourteen minutes.

The procedure was correct for the only consumer that existed while it ran. What made the gap permanent was the Monday, when the watcher went back out with four characters changed and those four characters were a group of its own. Galewatch acquired a second reader of that stream that morning, starting at the end of the log, and nothing anywhere recorded that fourteen minutes of July had just become unexamined by it.

The move that follows is small, and I have never seen a team make it before they needed to. Before you replay anything, or change an event's shape, ask the broker to list the consumer groups on that topic. The broker knows. Your diagram does not, and lesson 030's twelve component page shows how fast one of those goes stale at a four person shop. That list is the only honest inventory of who depends on you, and it is out of date the next time somebody deploys.

Publish the facts, call the decisions

Stagefront is this course's fictional ticketing service, where a stadium show goes on sale at exactly ten in the morning and two hundred thousand people press the same button in the same minute. Oversold seats are a lawsuit.

Everything in that sentence argues for events, and the seat claim cannot be one.

Lesson 015 settled the mechanism: clicking a seat commits a hold in one short transaction that talks to nothing outside the database, and the buyer then spends a minute and a half finding their card with no transaction open and no row locked, because the claim is already data. Lesson 024 established that Stagefront refuses to distribute that row at all, and pays for it with the ceiling lesson 013 declined to remove.

Why an event cannot do that job is the test. An event says something has already happened. A seat claim is a decision about whether something may happen. You cannot publish a decision you have not made. Two buyers publishing "I want seat 14F" is two statements of desire, something still has to pick, and the thing that picks is a row with a constraint on it.

Everything around the claim should be events: the confirmation, the reminder the night before, the PDF lesson 016 posed as a question, and lesson 003's admin reporting service, which has always been off the purchase path. Every hop you add to the path that takes money multiplies into the 19.3 hours a year 003's chain already costs.

The last of those is an event stream already and nobody calls it one. A replication stream is an event stream with one schema and one consumer, which is why lesson 012 could sit and watch Marlow's replica replay positions in a write ahead log. Lesson 011 is where Stagefront is told to point that reporting service at a replica of its own, and nobody there has done it. The pattern is older than the vocabulary, and buying the pattern does not come with the words.

Here is the line to draw, and it is the most useful sentence I know on this subject. Publish the facts, call the decisions. Marlow draws it at the card, because 017 would not put a waiting customer behind a queue. Stagefront draws it at the seat row. Both publish everything on the far side of that line, and both are right.

Nothing has a call stack any more

Lesson 026 printed an opinion and I want to go back to it, because April came at it from a side it had not considered. Marlow has two boxes, one database and a queue, 026 said, so it does not need a tracing system and should not install one this year; the full apparatus earns its money once a page is assembled from five services.

April added no services. It added two consumers, and 026's test never fires for that.

A trace in 026's terms is one request's path with a timing for each step, stitched together because every hop passed along the same id, and 026's own verdict is that the valuable part is the id, not the waterfall picture. With one consumer the shop already had an id that worked: the order id sat in the checkout's log line and the email consumer's, and that was the entire system. With three consumers the question stopped being how long this took and became what else happened because of this order, and the order id answers that only if all three of them write it down.

So the cheap fix is not a tracing system and never was. The message carries an id of its own and every consumer logs it: one field and an afternoon, which fits this course's record that its cheapest observability has been a table.

There is a standard shape for that field, and knowing it saves an argument. CloudEvents is a specification for the envelope around an event. It requires four attributes on every one: an id, a source naming who emitted it, a type and a specversion, with time and subject optional. Three of those four answer who said this, what kind of thing is it, and have I seen this one before, which is lesson 019's idempotency key arriving as a standard field instead of as something each team invents.

The trace itself is harder across a broker, and this is where an instrument people assume they have turns out not to be there. The web standard carries context in an HTTP header called traceparent, holding a version, a trace id, a span id and a flag byte. A message is not an HTTP request. Somebody has to copy that into the message's own headers and read it back on the other side; decent client libraries do it for you now, and it is the kind of thing nobody notices is missing, because nothing fails when it is.

Then the part that survives even perfect plumbing. It is not a waterfall. The producer's span ends when the publish call returns, which at Marlow is before the customer's page has finished loading. The consumer's work starts later: a second later, or thirty four seconds later on the March Thursday, or a month later on the third Tuesday in April. Nothing stops you recording the consumer's work as a child of that span, and the picture stops meaning anything the moment you do, because a parent's duration no longer covers what hangs off it. So the convention is a link instead: the consumer's work is its own trace, pointing back at the producer's. What you get is a forest, and reading one is a skill.

Lesson 026's limit gets worse here rather than better. Sampling keeps one request in a hundred, so a decision taken at the edge about one request now also decides whether you can see the three pieces of work it caused. 026 did this arithmetic on the right example: lesson 019's seventeen December card charges with no order behind them, at one percent sampling, is an expected zero point one seven traces. You will never see them. A rare path still needs a row rather than a sample, and the architecture changing has not changed that.

One last thing about the counter, where the morning started. Lesson 019 sorts work into three piles and prints a warning inside the first one: work that is idempotent for free should still be written down as such, "because the next person to add a counter to that handler will quietly move it into another one." April added the counter as a consumer rather than a line in a handler, which does the same thing to a whole system instead of one function. What the founder shipped on the Wednesday was not a deduplication table. order_count was deleted, and the graph now runs one query a minute against the orders table, which is the four minutes lesson 026 asked for in August, after eight months of looking for a reason to spend them. A number you recompute cannot be double counted, which puts it in 019's free pile and keeps it there. If a consumer's job is to count, count the thing the events are about rather than the events.

Recap

An event is a statement that something happened; a command is an instruction to one recipient. They travel through identical plumbing, which is why a queue full of commands can be mistaken for an architecture for years. The test is cheap: if adding a consumer forces you to change the producer, what you had was a command.

A producer that publishes a fact stops knowing who is listening. That ignorance is the entire product and the entire bill. Marlow bought two consumers for twelve lines each and never touched the checkout; what it sold in exchange was any ability to answer what a replay would do.

Each new consumer brings its own pile. Lesson 036 proved a broker cannot change which of lesson 019's three piles your work sits in. Fan-out multiplies the number of places you have to have got that right, and Marlow's April is the small version: three consumers, and only one of them had ever been asked whether a second delivery was survivable.

An event is a fact about the past and your database is the present. A consumer that joins the two is reading two different moments, and at Marlow the event is behind the shelf by a payment provider, which is 300 milliseconds on a good evening and past four minutes when lesson 019's sweeper has to clean up.

If the events do not cover every change to a thing, you cannot derive that thing from the events. The refund path publishes nothing, so the shelf is not reconstructable from the stream, and the first consumers to try will appear to succeed.

The broker's list of consumer groups is the only honest inventory of who depends on you. Galewatch replayed twenty one minutes correctly for the group that existed on the Thursday, and never at all for the group that appeared on the Monday. Ask the broker before you replay or reshape anything, and ask again after anybody deploys; the diagram is always older than the deploy.

Publish the facts, call the decisions. You cannot publish a decision you have not made, which is why Stagefront's seat claim stays a committed row and everything around it is events, and why Marlow's card call stays in the request where the customer is waiting.

Check your understanding

  1. Your team publishes user_signed_up and has three consumers. An engineer is about to replay last week's events and notices that they do not know what the third consumer does. Say what you would want to be true of the system before that replay is survivable, and which consumer you would check first.

  2. A colleague proposes putting the full order, line items and customer address included, into every order_placed event so consumers never query the database. Argue both sides with this lesson's figures, then say what you would need to know about the consumers before deciding.

  3. Marlow's reorder emails are useless because they fire per order against a shelf that has already moved. Redesign the feature without removing the event, and say which of lesson 019's three piles your design puts the work in.

  4. You inherit a service that publishes six event types. Nobody knows who consumes them. Lay out, in order, how you would find out, and say what you would do about an event type that turns out to have no consumers at all.

  5. An event-driven flow across four services is losing one order in a thousand and nobody can say where. You have logs everywhere and no tracing. Say what you would add first, what second, and which earlier lesson each belongs to.

  6. Somebody at Galewatch proposes folding the alarm check back into the six writer processes, on the grounds that one group reading the stream once is half the handler runs of two. Price what that saves, say what it gives up, and say which part of lesson 037's July Thursday the proposal is rebuilding deliberately.

Next lesson

039 Sagas and Distributed Transactions. Today's architecture publishes facts and lets anybody act on them; next lesson takes the case where several of those actions have to succeed or fail together across services that do not share a database, and prices what you give up to get it.

Finished reading?

Marking a lesson done keeps your place on the course index. It is stored only in this browser.

Tip: use the ← and → keys to move between lessons.