Reliability

Thirteen runs. Thirteen passes. Zero yen out of place.

Speed is the smaller half of a payment benchmark. Each run here generates load, most of them break something on purpose, and all of them end with a verifier that checks the outcome against what every client actually asked for.

150/ssustained with no failed requests
44 msmedian payment at that rate
63 ms99th percentile at that rate
13 / 13runs verified correct

What "pass" means

A run passes only if all four of these hold.

1. Every intent has one outcome

Each request a client made is replayed under its original key and must return what the client saw. Requests that never got an answer are replayed until they have one.

2. Every wallet adds up

Each wallet must hold exactly what those outcomes sum to, with nothing left reserved.

3. The ledger balances

The ledger's own invariant check must find nothing.

4. Settlement agrees

The day is settled, every shop must be paid exactly its net, and the ledger, the settlement records and the bank must match.

Latency when nothing is refused

The five runs in which no request failed, including three with a dependency taken away. Median and 99th-percentile payment latency, in milliseconds.

Payment latency by run Steady load: median 44 ms, 99th percentile 63 ms. Kafka frozen: 48 and 80. Hot merchant: 56 and 173. Bank frozen: 50 and 191. Merchant service killed: 47 and 366. 0100200300400 ms Steady, 150/s 6344 Kafka frozen 25 s 8048 Half to one shop 17356 Bank frozen 12 s 19150 Merchant service killed 36647 median 99th percentile

Every run

"Requests failed" counts attempts that got no response or a server error; clients retry them under the same key. "Unresolved at the end" counts intents whose client ran out of retries, which the verifier then resolves to a single outcome each.

RunWhat is done to itResultRateRequests failedp50 / p99 msUnresolved at the end
SteadyMixed traffic at a constant ratePASS149.9/s044 / 630
PeakTwice the steady rate, past saturationPASS178.7/s92.7%355 / 2,9234,092
Hot merchantHalf of all payments go to one shopPASS149.9/s056 / 1730
Retry stormClients give up after 40 ms and retryPASS99.9/s36.7%4 / 390
Ledger killedThe ledger is killed twice under loadPASS99.9/s38.4%252 / 1,7274,005
Payment service killedKilled twice under loadPASS82.2/s49.5%54 / 2,8161,080
Merchant service killedDown for 15 s; payments continue from cachePASS99.9/s047 / 3660
Kafka frozenFor 25 s; events must arrive afterwardsPASS99.8/s048 / 800
DynamoDB frozenFor 12 s; money-moving requests are refusedPASS100.0/s60.9%48 / 2,378773
Redis frozenFor 12 s; QR payments stop, the rest continuesPASS92.3/s044 / 1,4110
Bank frozenFor 12 s; top-ups wait, then complete oncePASS100.0/s050 / 1910
Database frozenTiDB for 8 s; requests stall and resolvePASS99.9/s7.0%46 / 2,327636
EverythingSix faults in 90 s: two kills of the ledger, one of the payment service, Kafka, the bank and Redis frozenPASS92.4/s40.6%248 / 2,6345,026

Is the verifier any good?

A check that has never failed proves nothing. So a bug was planted: the payment service was rebuilt to ignore idempotency keys, and the retry storm was run against it.

200 of 200 wallets wrong

Retried top-ups and transfers were applied more than once, and the verifier caught every one.

10 of 10 shops mispaid

Each was paid something other than its true net.

The ledger still balanced

Every duplicate was a perfectly balanced, perfectly reconciled movement of money that nobody had asked for. Only comparing outcomes with intents finds that kind of fault.

How to read these numbers

These are single-machine numbers. Every service and every store ran in containers on one modest computer, one instance of each. They are useful relative to each other, not as a forecast of production throughput. With a single instance, a killed service is simply gone until it restarts, so the failure drills measure correctness under failure and not availability. Past roughly 150 to 300 requests a second on that machine most requests time out, and it stays correct.

Each run's full report is written by the tool that ran it, and all thirteen can be reproduced in about twenty minutes.