Marlow Books is an online bookshop run by four people, and it exists only in this course. On a Thursday evening in March, a month after the founder's week off in lesson 035, it sent one customer three identical order confirmations and then filed that customer's order in the drawer it keeps for work that failed.
Nothing failed. That is the whole of today.
The email provider went slow at twenty to eight. Not down, slow: it accepted every message and took between thirty one and forty seconds to say so, against a median the founder measured the following week at 240 milliseconds. Lesson 019 has the sentence for this and it is the one to carry into today. A dependency that is down is safe. A dependency that is slow is the one that duplicates.
Here is the piece of lesson 018 that decides what happens next, published without a name on it. When the email consumer tells the broker a message failed, 018 says, "the broker offers that message to somebody else in thirty seconds and hands out other work meanwhile." Thirty seconds is doing more work than a retry interval. It is the broker's hold on a message it has handed out: how long it waits for an answer before deciding the holder is gone. Lesson 018 printed the number on the path where the consumer speaks up, and that the same timer runs when nobody does is my reading of its sentence rather than a line in it. A message is never removed when it is delivered. It is hidden, and the hiding has an expiry.
So take one order placed at 19:52:00, and give the handler the normal shape: pick up the message, call the provider, and only when the provider answers tell the broker you are done.
At 19:52:00 the broker hands the message to a worker on box B, which calls the provider. The email leaves. At 19:52:30 the provider has still not answered, the thirty seconds are up, and the broker hands the same message to a worker on box A. The email leaves again. At 19:52:34 box B's call returns, box B tells the broker it finished, and the broker refuses, because the hold box B is quoting expired four seconds ago. Some brokers do the opposite and accept it, which deletes a message another worker is holding right now, and that is lesson 035's DEL removing somebody else's lock arriving in a queue. At 19:53:00 the message goes out a third time. On the Friday of Christmas week in lesson 018 the founder set a limit of three: a message handed out three times moves aside instead of coming back. At 19:53:30 that limit was reached and the broker moved this one to the dead letter queue, thirty seconds after the third copy of the email had already left the building.
Three emails, one order, one charge, and a message filed under failure with every bit of its work done.
That ran from 19:40 to 21:10. At lesson 027's figure of one order every fifty seconds, which carries lesson 009's Christmas conversion rate onto an ordinary evening and which nobody at the shop has ever checked, ninety minutes is 108 orders. Three hundred and twenty four emails. A hundred and eight messages in a dead letter queue that nothing was watching, which is lesson 018's warning about that queue, that it needs an owner and an alert on being non-empty at all, coming true.
Lesson 017's age alert never fired, and the reason is worth more than the alert. Marlow runs four of these workers, two on each box, which at the Christmas rate of a fifth of a message a second and a handler that normally takes a quarter of a second is about eighty times more consumer than the shop needs. Even with every call taking thirty four seconds and every message going round three times, the four workers were about half busy and the oldest unprocessed message never got past about a minute and a half. Spare capacity is a good thing to have and it turned a ninety minute failure into a graph with nothing on it.
What found it was a person. One customer read the third copy and replied to ask whether they had been charged three times. They hadn't. The card is charged in the checkout and not by this consumer, and lesson 019's idempotency key has kept that one call honest since January. Three identical emails give a customer no way to know that. The founder spent Friday morning proving there had been one charge, found the dead letter queue on the way, and came out of it with a proposal that is the single most common fix for this bug and is worse than the bug.
Where you put the acknowledgement is the whole guarantee
An acknowledgement, or ack, is the consumer telling the broker that a message is dealt with and can be forgotten. Until it arrives the broker keeps the message, because a queue is lesson 017's buffer with a memory and the memory is the product.
You have exactly two places to put it, and your delivery guarantee is nothing more than which one you picked.
Ack after the work, which is what Marlow had. If the consumer dies between doing the work and acking, the broker still holds the message and gives it to somebody else, so the work happens again. That is at-least-once: every message is processed one or more times, and the gap you are exposed to is the time between the effect landing and the ack arriving.
Ack before the work. The broker forgets the message immediately, then you do the work. If the consumer dies in that gap, nothing holds the message, nobody retries, and the work never happens. That is at-most-once, and the gap is the same gap read from the other end.
The founder's proposal on the Friday was to move the ack above the provider call. It is a tidy one line change and it does cure the hook completely. There is no second delivery to cause a second email, because the message is gone from the broker before the provider is ever called.
Run lesson 018's Friday through it before you ship it. At ten to five on the Friday of Christmas week a new field went into the confirmation payload, producer and consumer together, and forty messages written by the old producer were still on the queue. The new consumer raised a KeyError on the first of them, exited with nothing catching it, and the supervisor fed the fresh process the same message. Nothing else moved. Lesson 017's age alert fired at ten minutes, the forty went to the dead letter queue, a five line script gave them their missing field, and they were back on the main queue before six.
Under ack-first, every one of those forty is acknowledged and then dropped. Take, ack, raise, die, restart, take the next. The queue drains in well under a minute and there is no head of line blocking at all, because the head of the line keeps being deleted. Nothing reaches the dead letter queue, because a dead letter queue is something a broker does to a message it still holds and it holds none of these. Lesson 017's age alert never fires, because the oldest message is always a second old. Forty customers never get a confirmation and nothing anywhere knows their names.
At any single instant the exposure is small: whatever the consumers have acked and not finished, which for four single threaded workers is four messages. Across an incident it is the whole incident, because the exposure refills as fast as the queue drains. And the loss arrives the way lesson 018 taught you to fear most. The absence of an event is not an event, so nothing in the building can see it, while a duplicate turns up in somebody's inbox the same minute.
| Where the ack goes | You get | What the gap costs | Marlow |
|---|---|---|---|
| After the work | at-least-once | duplicate effects | three emails |
| Before the work | at-most-once | lost work, silently | forty emails, never |
The third row does not exist. The work and the ack are two operations against two machines with a network in between, and nothing makes them atomic. Anyone selling you exactly-once delivery is selling you one of these two rows with something bolted on top.
My own default is at-least-once, always, with the duplicates handled deliberately. A duplicate is a bug you fix once, in code you own, with a test. A lost message is a bug you rediscover from customers, forever, and you never learn how many you missed.
The visibility timeout is a lease, and you have met it twice already
The hiding has names. Amazon SQS calls it a visibility timeout and defaults it to thirty seconds, which is where Marlow's number came from. Google Pub/Sub calls it an ack deadline. Kafka has no per message hold, because a consumer there tracks a position in a log rather than holding messages one at a time, which lesson 037 takes up tomorrow. Its equivalent is max.poll.interval.ms: miss it and the group hands your partitions to somebody who resumes from your last committed offset, which redelivers the same work by another route. The shape is the same wherever it exists: the broker starts a timer, and when the timer runs out without an ack somebody else gets the work.
Read that again as lesson 035 would. The broker cannot see your handler. It cannot see the email provider. All it has is silence, and it has to turn silence into a decision on a timer. A lock is a promise about a resource, and almost every implementation makes a promise about a client instead, which lesson 035 said about the lock services you would go out and buy, and which is just as true of a queue you did not write and probably did not configure. The visibility timeout is a lease, with lesson 032's expiry and lesson 034's clock underneath it, and the only new thing about it is that it came switched on with a default somebody else chose.
So a redelivery is not evidence that anything died. It is evidence that a timer expired. In the hook, nothing crashed, nothing errored, no process restarted, and every message was delivered the maximum number of times the broker permits.
The arithmetic you want is a ratio. Marlow's handler takes 240 milliseconds against a thirty second timeout, which is a margin of 125 times, and that sounds like a margin nobody needs to think about. It is a claim about the median. Lesson 020 says a timeout is a claim about somebody else's latency distribution; a visibility timeout is a claim about your own handler's, and almost nobody graphs that one. On the Thursday the ratio inverted to 34 against 30 and the duplicate rate went from zero to every single message, with no intermediate state to notice. A margin that only exists at the median is not a margin.
Two things you can do, neither free. Extend the hold while you work, which most client libraries will do for you on a background thread, and lesson 032 has already priced it: a renewal thread cannot tell whether the work is moving, so it will cheerfully extend a hold around a wedged handler, and you have bought a message nobody will ever retry. Or raise the timeout, and a dead consumer's messages sit invisible for however long you set, which becomes your worst case recovery time.
Then there is the drawer. Lesson 018 defines the dead letter queue as "a second queue that the broker moves a message onto once it has failed too many times or run past its deadline", and both halves of that sentence are right. The second half is the one nobody watches. A delivery count counts deliveries, not failures, so a dead letter queue holds the messages the broker gave up on, which is not the same set as the work that did not happen. Marlow's held 108 messages whose work had been done three times each. Before you write a script that replays a dead letter queue, work out which of those two sets you are about to replay.
What the vendors are actually selling
Several of them do sell exactly-once, the feature is real, and it is narrower than the name.
Start with the smallest version. SQS FIFO queues take a message deduplication id, and the broker ignores a second message carrying an id it has already seen within a five minute deduplication interval. That is a genuine guarantee, and the number in it is the useful part. A window is an admission: the vendor is telling you how long their memory is, and every duplicate later than that is yours. Lesson 019 published the rule for comparing the two sides and it transfers without a change: your key's lifetime has to outlive your own retry deadline. A producer that backs off one minute, then two, then four has its fourth attempt at minute seven, two minutes past the end of the broker's memory, and the careful retry policy is what defeated the deduplication.
Kafka's version is better and still has edges. With enable.idempotence on, which has been the default since Kafka 3.0, the producer asks the broker for a producer id, then stamps every batch it sends to a partition with a sequence number that only goes up. The broker remembers the last sequence it accepted for that producer and that partition, and silently drops a repeat. That is lesson 034's rule implemented by somebody else: the identity of an event should come from a counter, not from a clock, and a counter the writer owns is immune to everything that makes a timestamp a measurement.
What it covers is the producer's own retries, which is lesson 020's duplicate: the one a client creates when it does not get an answer. It does not know your application called send twice for the same order, because to the producer those are two different events. And without a configured transactional.id, a restarted producer asks for a new producer id, so the broker has never seen it and the memory starts empty.
Give it a transactional.id and you get the real thing. The producer can open a transaction, write records to several partitions, and commit the consumer offsets it is advancing into that same transaction, because in Kafka the offsets are themselves just another topic. A consumer set to isolation.level=read_committed never sees records from a transaction that is still open or was aborted. Read a topic, transform, write a topic, commit the offset, all or nothing. That is exactly-once processing and it is not marketing.
It holds while every single thing the job touches is Kafka.
Point the output at an email provider and the transaction does not stretch to cover it. Point it at Postgres and it does not stretch there either, because Kafka's transaction and Postgres's transaction are two protocols that have never been introduced. A broker can only make exactly once true about the things it owns. Your effect is in your database, or in somebody's inbox, or on a card, and none of those belong to the broker. Lesson 040 is the pattern for the Postgres half and it is called the outbox.
Two phrases get used as synonyms and they should not be. Exactly-once delivery is a claim about how many times a message arrives, and that is the lie, because the network cannot tell a lost message from a lost acknowledgement. Exactly-once processing is a claim about how many times the effect happens, and that one you can build.
The effect is the only place you can make it once
Lesson 019 sorted work into three piles, and the piles decide this entirely. What the broker is does not.
If the effect is a row in your own database, you are in the second pile and the whole problem costs five lines. Put the deduplication and the effect in one transaction:
begin;
insert into handled_messages (message_id) values ($1)
on conflict do nothing
returning message_id; -- no row back means somebody did this already
-- only if a row came back:
update orders set status = 'paid' where id = $2;
commit;
Carried out in words, because the audio skips the block: you insert the message id into a table of ids you have already handled, with on conflict do nothing so a repeat is refused rather than raised, and returning so the handler can tell which happened. No row back means stop, commit nothing, ack. The effect and the record of the effect commit together or not at all, because they live in the same database, and that is the only reason it works. At-least-once delivery plus one unique constraint gives you exactly-once processing with no vendor in it and no window, right up to the day somebody puts a retention policy on that table, which is lesson 019's expiry question in a different schema.
Galewatch, which collects a reading from each of its wind turbines every two seconds, has the cheapest version of this in the course, and there is no broker in it anywhere. There are 1,240 turbines at lesson 034's count, so 620 readings a second. A turbine whose write times out sends the reading again a minute later, which is at-least-once implemented in hardware on a hillside, and lesson 019's unique constraint on the reading's own identity throws the second copy away. The data names itself, so the deduplication is a primary key doing its ordinary job at 620 a second with no coordination of any kind.
And it still has a window, which is the part worth keeping. Lesson 034 found it: the constraint's idea of identity is a turbine and a clock reading, so a daemon that steps its clock back six seconds deletes three genuine readings instead of three duplicates. Every deduplicator has a window. SQS FIFO's is five minutes, Kafka's is however long the broker keeps producer state, and Galewatch's is the accuracy of a clock on a pole in the rain.
Stagefront, the ticketing service where two hundred thousand people press the same button at ten in the morning, is the case where a duplicate costs real money, and it needs nothing from a broker either. A seat is claimed with update holds set status = 'sold' where status = 'held', which lesson 019 puts in the first pile and lesson 035 calls the row's own status working as a fence with two values. The second delivery runs the same statement, matches zero rows, and the consumer treats zero rows as success, because that is what zero rows means here. No lock, no token, no deduplication table.
Which leaves Marlow, in the third pile, where you cannot make the provider forget. You can still make the decision to send idempotent, and you pay for it in a currency this lesson has already named. What the shop shipped the week after the Thursday is a pattern it already owned. Lesson 019 stamps a sent_at on a row per publisher before the payout emails go out, and nobody had carried that across to confirmations. Now they have it too: a row claimed before the send, carrying the order id and a null sent_at, stamped once the provider answers. A second delivery finds that row and has to decide what a null sent_at means, which is lesson 019's unanswerable question wearing a hat: the first attempt is either still in flight or dead. Marlow's rule is that a claim younger than two minutes is in flight, so the second delivery hands the message back to the broker with a two minute delay instead of acking it, and anything older gets sent.
Run the Thursday through it. The provider's worst answer was forty seconds, so the redelivery at thirty seconds always finds a claim half a minute old, defers, and comes back at two and a half minutes to a stamped sent_at and an easy ack. Zero duplicates, zero losses, nothing in the drawer. Now break it: kill the worker a second after the claim, and that same delivery finds a claim older than two minutes, decides the first attempt is dead, and sends. The customer gets their confirmation two and a half minutes late. That was delivery three, the last one the broker will make, so the recovery fits with nothing to spare, and nobody at the shop has noticed that a two minute window, a thirty second hold and a limit of three leave no room for a fourth try.
The residual is a provider that genuinely takes over two minutes to answer, which sends a second email, and Marlow has written down that it would rather send two than none. Its two minutes is a window chosen by Marlow rather than by a vendor, and it is the same admission. The broker does not change which pile your work is in, and the pile is what decides the bill. Swapping SQS for Kafka moves nothing in the four paragraphs above.
The promise you did not buy
People reach for a queue and quietly assume it preserves order. Be precise about what is on offer.
Inside one Kafka partition there is a real order, and it is real for the reason lesson 034 gave about Marlow's single Postgres: the log decides an order rather than measuring one, so there is no clock in the answer and nothing to be wrong about. Across partitions there is none. A consumer group reading six partitions is six readers who never speak, and lesson 034 named that exactly: a logical clock orders what talks and has nothing to say about two things that never spoke.
At-least-once fights order too, in the same mechanism that creates the duplicates. A redelivered message arrives after the messages that came behind it. Lesson 018 put it as backoff punishing the oldest work, and a dead letter queue is the extreme version, where the message leaves the order entirely and comes back whenever somebody runs the replay script.
So pick the key whose order you actually need, partition by that, and write the choice down. Marlow needs none: confirmations for different orders are independent, and two for the same order is today's bug rather than an ordering problem. Galewatch does not need arrival order either, since every reading carries its own stamp, with lesson 034's warning attached about what that stamp is worth. Stagefront's order is per seat and the database settles it.
The price of ordering by key is that one key is one consumer, and your throughput ceiling becomes the busiest key, which is where lesson 047 picks it up. Order everything globally and you have one partition, one consumer and lesson 001's single server with better marketing.
Recap
Where you put the acknowledgement is the whole guarantee. Ack after the work and a failure in the gap repeats it. Ack before and a failure in the gap loses it. There is no third placement, because the work and the ack happen on two machines and nothing makes them atomic.
A redelivery is not evidence of a crash. It is evidence that a timer expired. The visibility timeout is lesson 032's lease with lesson 034's clock under it, and its safety margin is usually a claim about your handler's median that nobody has graphed.
A dead letter queue holds the messages the broker gave up on, which is not the same set as the work that did not happen. Marlow's held 108 messages that had been fully processed three times each. Check which set you have before you replay it.
Every deduplicator has a window, and the window is an admission. Five minutes in SQS FIFO, whatever the broker keeps of a producer's state in Kafka, the accuracy of a turbine's clock at Galewatch, two minutes in Marlow's own sent_at rule. Lesson 019's test applies to all of them: your key has to outlive your own retry deadline.
A broker can only make exactly once true about the things it owns. Kafka's transactions are not a lie and they stop at Kafka's edge, because an offset is a Kafka record and your effect is not. The moment the effect leaves, you have two systems and no transaction across them.
The broker does not change which pile your work is in, and the pile is what decides the bill. A database write costs one unique constraint inside the effect's own transaction. A third party costs you a choice between sending twice and sending never, and the honest move is to pick on purpose and write the choice down.
Check your understanding
Take the hook and change one thing: Marlow's broker accepts a late acknowledgement and deletes the message, instead of refusing it. Walk the 19:52 timeline again, say how many emails that customer gets, and say which of the two broker behaviours you would rather have and why.
A colleague moves the ack above the work for a consumer that writes a row to Postgres and nothing else, arguing that duplicates are the real risk. Make their strongest case, then say what you would do instead and roughly how much code it is.
Your team buys a queue whose documentation promises exactly-once delivery. Write the three questions you would ask before believing it, and say what answer to each would make you change your handler anyway.
A payments service consumes a topic and calls a card processor. It has Kafka transactions turned on end to end and the architect says duplicates are impossible. Say precisely where that is wrong, and design the smallest fix that does not involve changing brokers.
Galewatch's deduplication is a unique constraint and its window is a clock's accuracy. Stagefront's is a status column and appears to have no window at all. Say whether that is true, and if it is, say what Stagefront has that Galewatch does not.
You inherit a dead letter queue with 40,000 messages in it and a replay script somebody wrote and never ran. Say what you would measure before running it, and name the one property of the handler that decides whether replaying is safe at all.
Next lesson
037 Event Streams: Kafka and the Log as a Database. Today treated the broker as something that holds a message until you are done with it; next lesson takes the other model seriously, where nothing is held or deleted and a consumer is just a position in a log that everybody can read.