Picture this. You walk into a room (or a Zoom call). A stranger smiles, says hi, and gives you one sentence: "Design Instagram." Then they lean back and wait.
Forty-five minutes. One vague sentence. And on the other side of that conversation: a level, an offer, maybe a life-changing amount of money. The strange part? The interviewer is not really testing whether you can design Instagram. Nobody builds Instagram in 45 minutes. They're testing something else entirely — and most candidates never figure out what.
This chapter is the answer key. Before you learn a single thing about caching or sharding, you're going to learn what this exam actually grades. Study the exam before the subject — it changes how you learn everything that follows.
The interviewer sits there with a mental scorecard, not a checklist of technologies. Across companies the scorecard looks remarkably similar. They're asking themselves:
A system design interview is a driving test, not a written exam. The examiner doesn't care whether you can recite the manual — they care whether you check the mirror before changing lanes, without being told. Naming ten databases is reciting the manual. Saying "I'll put a queue here, because this write path will spike at 9pm and I don't want the database taking that punch directly" — that's checking the mirror.
Nearly every strong interview follows the same skeleton. Learn it like a song structure — verse, chorus, verse. You'll practice it so many times in this book that it becomes muscle memory.
"Design Instagram" is a trap if you take it literally. Instagram is fifty products in a trench coat. Your first job is to shrink it, out loud, with the interviewer's agreement: "I'll focus on photo upload and the home feed. Stories, DMs, and Reels are out of scope — fair?"
Split what remains into two lists. Functional requirements — what the system does: upload a photo, follow people, see a feed. Non-functional requirements — the qualities it must have: feed loads in under 200 milliseconds, survives a machine dying, handles 500 million users. Non-functional requirements are where design decisions come from. A feed for 500 users and a feed for 500 million users are different machines that happen to share a name.
Back-of-envelope math has one purpose, and it isn't impressing anyone with arithmetic. The numbers must drive decisions. If your estimate doesn't change what you build, skip it — that's the senior move.
"500 million users, maybe 100 million daily. Each opens the feed a few times a day — call it 5. That's 500 million feed loads a day, roughly 6,000 per second on average, so plan for 20,000 at peak. Writes? Maybe 100 million new photos a day, about 1,200 per second. So this system is enormously read-heavy — reads outnumber writes maybe 20 to 1. That ratio is why I'll design the read path first, and why caching will carry this whole system."
See what happened? The math produced a conclusion, and the conclusion shaped the architecture. That's the whole game. Chapter 1.4 drills this until it's reflex.
A handful of endpoints and the core tables/entities. POST /photos, GET /feed?cursor=…. A users table, a photos table, a follows table. Don't gold-plate this part — its job is to prove you can turn requirements into concrete shapes, then get out of the way.
Boxes and arrows: clients, load balancer, app servers, database, cache, queue, storage. Here is the single most important rule at every level: finish the loop. A request must be traceable from the user's thumb all the way in (write path) and all the way back out (read path). The most common failure in these interviews — documented over thousands of real sessions — isn't a lack of depth. It's never finishing a working design. An incomplete brilliant design loses to a complete boring one, every time.
Everything before this was the qualifying lap. Now the interviewer picks a box on your whiteboard and says "tell me more" — or, if you're playing at the senior level, you pick the box before they ask. This is the part of the interview that separates levels, and it deserves its own section.
In every design problem, most of the system is routine — CRUD endpoints, a load balancer, a database. And then there's one part that's actually hard, the part the question was invented to probe. For a URL shortener, it's ID generation without collisions. For a chat app, it's message ordering and delivery guarantees. For a news feed, it's fanning out posts from accounts with 100 million followers. Finding the crux fast — and saying it out loud: "the interesting problem here is X, I want to spend most of our time there" — is the single strongest senior signal that exists. Every problem chapter in Part 3 stamps its crux exactly like this box, until spotting them becomes instinct.
A design problem is like a house tour where one room is on fire. Junior candidates tour every room with equal enthusiasm, describing the curtains. Senior candidates smell smoke, walk straight to the burning room, and start talking about how to put it out. The interviewer already knows which room is on fire — they lit it. They're waiting to see how long you take to notice.
Here's the uncomfortable truth from people who've sat on both sides of the table: a complete, sensible high-level design with reasonable trade-offs is a strong pass for a mid-level candidate — and the same performance gets marked "lacking depth" for a senior. The delta isn't knowledge. It's behavior.
| Dimension | Mid-level (E4 / L4) | Senior (E5 / L5) |
|---|---|---|
| Who drives | Covers requirements and the high-level design, then lets the interviewer steer the deep dives | Drives all 45 minutes; proposes their own deep-dive agenda: "I want to dig into the fanout path next" |
| Trade-offs | Discusses options when asked | Commits with justification: "Postgres, because I need transactions here and it won't be the bottleneck" — then names the cost of that choice |
| Depth | Solid breadth; textbook understanding is acceptable | Goes deep on 2–3 components without being asked, with specifics: what's cached, what happens on a miss, what happens when the cache dies |
| Failure | Handles failures when prompted | Raises them first: "What happens at 10x load? When this region dies? Who gets paged?" |
| Numbers | Forgiven for rough estimation | Estimation is non-negotiable, and each number must earn its place by changing the design |
Read the second column again. Nothing in it requires knowing more facts. It requires owning the room: proposing, committing, self-critiquing. That's trainable, and it's precisely what the Level 3 quizzes in every chapter train — they put you in a scenario and make you choose like an owner, then explain why hedging reads junior.
A candidate interviewing for a senior role at Meta described their design round like this: the interviewer kept asking one question, over and over — "and what if you had 100 times more users?" Each time, part of the design melted. A single database became sharded. A simple fanout hit a celebrity account and had to become a hybrid push-pull system. The candidate redesigned their sharding strategy three times in one interview. That wasn't the interview going badly — that was the interview. Meta's design rounds deliberately push a working design until it breaks, because how you respond to your own design breaking is the senior signal. Expect it, and you can even enjoy it: when your design breaks, say "good — here's what breaks first, and here's the fix," and you're demonstrating exactly what they came to see.
This is the newest and least-prepared-for part of the senior bar. Interviewers increasingly grade whether you think past the launch: How do you know the system is healthy? What do you watch — latency, error rate, queue depth? How do you roll out a change without betting the company on it? How do you roll back? What does the 3am page look like, and what did you build so it never fires?
One retelling from a failed senior loop puts it perfectly: "You never said monitoring, deployment, or rollback. You never walked through what happens at 3am. That gap alone can sink an otherwise strong performance." The fix costs thirty seconds: end every design with what you'd monitor and how you'd ship it. We'll build this habit in Part 2 (observability, rollouts) until it's automatic.
Thousands of post-mortems from real loops boil down to a short list. Pin it somewhere.
Jackson Gabbard spent years at Facebook and conducted hundreds of interviews there. In 2016 he did something interviewers rarely do: he wrote publicly about what he was actually grading in the architecture round, and the post has been circulating among candidates ever since.
Two lines from it are worth memorizing. The first: your level comes primarily from your performance in the design and architecture interview. Not from the coding rounds — those mostly prove you can code. The design round is where the compensation band gets chosen.
The second is sharper, and it stings: if you're looking to your interviewer for guidance about how to proceed, you're losing points. Not "you're neutral" — losing. Every glance at the interviewer that means is this okay? should I keep going? is a small deduction, because senior engineers don't need permission to work on a problem.
He also described what the strong candidates did instead. They turned a fuzzy requirement into concrete numbers using informed estimation — not uniform arithmetic, but realistic usage: users cluster in cities, traffic clusters at 9pm, most accounts are quiet and a few are enormous. They made trade-offs intentionally and said so. They drove the session. And they played to their strengths while openly flagging the things they didn't know, which — counterintuitively — read as confidence rather than weakness. Nobody expects you to know everything. They expect you to know which things you know.
How this chapter's material actually surfaces in a real round — the meta-questions behind every question:
Explain to a junior engineer, in five or six sentences, what a system design interview actually measures — and why "knowing lots of technologies" isn't it.
It measures judgment, not vocabulary. The interviewer gives you a vague problem and watches whether you can make it concrete: narrow the scope, put numbers on the load, and design a complete machine that a request can actually travel through. Then they watch what you do with the hard part — every problem has one genuinely difficult piece, and finding it yourself is worth more than knowing extra tools. Committing to choices and explaining why beats listing every alternative, because that's what owning a real system looks like. And they listen for signs you've lived with software in production: what you'd monitor, what breaks under load, what the failure plan is. A candidate with fewer facts and real judgment beats a walking encyclopedia every time.
Take "Design WhatsApp" and do only the first five minutes: write functional requirements (pick 3–4, cut the rest out loud) and non-functional requirements (with numbers you invent but justify). Time yourself. Five minutes, no more.
One good shape: functional — 1:1 messages, delivery/read receipts, online status; out of scope — groups, calls, stories, encryption details. Non-functional — messages delivered in under a second when both users are online; no message ever lost once the server accepts it (that acceptance acknowledgment is a promise); 1 billion users, 50 billion messages/day ≈ 600K messages/sec average, plan for 1.5M/sec peak; chat history available on a new device. Notice how each non-functional line implies architecture: "never lost" means durable storage before acknowledging, "under a second" means persistent connections, not polling.
For each system, name the crux in one sentence: (a) a URL shortener, (b) a ticket-booking site the day Coldplay tickets drop, (c) a Google Docs clone.
(a) Generating billions of short, unique IDs fast, without two servers ever producing the same one. (b) Thousands of people buying the same seat in the same second — concurrency and overselling, not traffic volume, is the fire. (c) Two people typing in the same paragraph at the same time — merging concurrent edits so every copy converges to the same result. If you found these in under a minute each, you're already thinking like the interviewer.
Rewrite this option-listing answer as a committing answer: "For the feed we could use push, where we write to every follower's feed at post time, or pull, where we build the feed at read time. Push is faster to read but expensive for popular accounts. Pull is cheaper to write but slower to read. Both are valid."
Something like: "I'll use push — write each new post into followers' cached feeds at post time — because feeds are read 20 times more often than they're written, so I want reads to be a cheap cache hit. The cost is that a celebrity with 100M followers triggers 100M writes, so for accounts above roughly a million followers I'll flip to pull and merge their posts in at read time. Push by default, pull for celebrities." Same facts, plus a decision, a reason tied to the read/write ratio, an honest cost, and a mitigation. That's the difference between a book report and a design.
Rohan and Priya prepped together for three months. Same books, same mock interviews, same 45-minute skeleton from the last chapter. Rohan interviewed at Google and walked out with an L5 offer. Priya took the identical preparation to Palantir a week later — and got quietly dismantled in a round neither of them had ever heard of. No load balancers. No scale numbers. Just an interviewer saying: "An infection is spreading through a city. Help us track it." Priya started drawing boxes and arrows. The interviewer let her finish, then asked one question: "Who looks at this system, and what do they do with what they see?" She had no answer. She'd designed a machine for a problem nobody had defined.
Meanwhile their friend Arjun, interviewing at Anthropic, hit a different wall. His design round wasn't "design Instagram" at all. It was "design the serving system behind a chat assistant" — and within ten minutes the interviewer was asking about GPU memory and what happens to latency when someone pastes a 200-page document into the chat. Arjun knew feeds and fanout cold. He'd never once thought about what a language model costs to run.
The skill you started building in chapter 1 transfers everywhere. The exam does not. Each of these four companies has spent years evolving its design round into a distinct instrument, tuned to detect a specific kind of engineer. Walk into the wrong exam with the right skill and generic prep, and you can still lose.
This chapter is the field guide: what each loop looks like in 2024–2026, what each one grades hardest, and how to point your preparation at the exam you'll actually sit.
Everything from chapter 1 still holds — scope the problem, find the crux, commit with reasons, finish the loop, end with operations. Think of that as your fitness base. What changes per company is which muscle gets tested to failure. Meta tests speed and iteration under pressure. Google tests depth and arithmetic. Palantir tests whether you can find the problem at all. Anthropic tests whether you can design for a workload — large language model serving — that most engineers have never operated.
Four exam boards, one subject. Board one runs a rapid-fire viva: confident answers, numbers from memory, and the examiner interrupts to change the question mid-answer (Meta). Board two wants a long written derivation — full working shown, every figure justified (Google). Board three hands you a word problem with no numbers in it at all and grades whether you ask the right questions before computing anything (Palantir). Board four is a lab practical on equipment you have to learn before exam day (Anthropic). Same physics. A student who only ever practiced one format walks into the others underprepared — not for lack of knowledge, but for lack of the format's reflexes. The analogy breaks in one place: unlike school, you get to choose which boards to sit, so choose with your eyes open.
A Meta E5 loop has multiple coding rounds and a behavioral round, but only one design round. One conversation, roughly 45 minutes, carrying enormous weight — because at Meta the design round is the primary leveler. Meta rarely rejects a strong candidate for a wobbly design round; it down-levels them. Strong coding, strong behavioral, unconvincing design is the classic recipe for an E4 offer when you interviewed for E5. Everything else being equal, this single round is often the difference in level, scope, and a very large amount of money.
Here's the part most candidates discover too late: there are two different design rounds, and a role tag chosen when you apply — not on interview day — decides which one you sit.
System Design is the backend-guts exam: design a distributed message queue, a web crawler, a top-K trending system, a rate limiter. Servers talking to servers. Product Architecture is the user-product exam: design the news feed, design Messenger, design Instagram Stories. It looks friendlier — everyone has used Messenger — but it grades things System Design barely touches: the client-server interaction (what exactly travels over the wire, and when), API design (pagination contracts, idempotency, versioning), the data model, and live updates (how a new message appears on your friend's phone in under a second — persistent connections, reconnect logic, what the client does when it comes back from a tunnel). Backend engineers who pick Product Architecture because "feed questions look easier" routinely get shredded on client details they've never operated.
Whichever flavor you sit, Meta grades four competencies: problem navigation (can you scope the vague prompt and find the hard part), solution design (a complete, reasonable architecture), technical excellence (at E5 this means real depth in at least one area — one component where you clearly know more than the diagram shows), and communication (can the interviewer follow you without effort).
Two style notes that define the Meta room. First, iteration is the format. Meta interviewers deliberately push a working design until it breaks — "and what if you had 100x the users?" — and grade how you evolve it. The chapter 1 war story of a candidate redesigning their sharding three times in one round was a Meta round; candidates report this experience again and again. Expect your design to break twice. That's the exam working.
Second, numbers and opinions are expected from memory. Meta rounds move fast, and stopping to hedge reads badly. "Why Redis over Memcached?" is a real, recurring probe — and the passing answer names a mechanism ("I need sorted sets and persistence here"), not popularity. The latency numbers below should be reflex, because they justify choices in one sentence: "that's a disk seek per request, ten milliseconds, so this must come from memory."
| Number | Rough value | The decision it drives |
|---|---|---|
| RAM access | ~100 ns | Cache hits are effectively free |
| SSD random read | ~100 µs | SSD-backed stores can serve reads directly |
| Spinning-disk seek | ~10 ms | Random reads on HDD cap you near ~200/sec/disk |
| Round trip, same datacenter | ~0.5 ms | A few internal hops per request are fine; twenty are not |
| Round trip, cross-continent | ~70–150 ms | Geo-replication can't hide in a 200 ms budget |
| One Redis node | ~100K ops/sec | One node carries a lot before you shard the cache |
| One relational DB node (writes) | a few thousand/sec, safe planning figure | Past this, you're talking queues, batching, or sharding |
The most consistently retold Meta outcome isn't a rejection — it's the down-level. The pattern shows up over and over in candidate reports from senior loops: coding rounds strong, behavioral strong, design round "fine but shallow" — complete architecture, no deep area, iteration handled hesitantly. The result is an offer, but at E4, with recruiters telling candidates more or less openly that the design round is what separates the two levels. One widely-echoed retelling describes finishing the loop confident, then hearing the recruiter say the design interviewer "didn't see E5 depth" — a phrase that maps exactly to the technical-excellence competency: at E5, breadth plus one genuinely deep area is the bar, and a round with no deep dive can't clear it no matter how tidy the boxes are. The practical lesson: at Meta, plan before the interview which component you'll take deep, because the level rides on that one round more than on any other.
Google's design prompt is vague on purpose, and often narrower than you expect: not "design Google Docs" but "design the commenting system in Google Docs." The narrowness is a gift — it points at where they want depth — but the vagueness is the first test. The interviewer will not run the meeting for you. Google interviewers are famously quiet: they prompt, take notes, and watch. Many candidates read the silence as disapproval and start flailing. It isn't disapproval. It's the format. You drive, or nobody does.
A strong Google 45 minutes has a distinctive shape — much less time on requirements than you'd guess, much more on depth:
The deep dive is explicitly what separates L5 from L4. An L4 performance completes a sensible design and answers follow-ups competently. An L5 performance picks two or three components — unprompted — and goes to the floor of them: what's stored, in what layout, what a single request costs, what breaks first, and the numbers behind each claim.
Which brings us to the most Google-specific thing in this chapter. NALSD — Non-Abstract Large System Design — is a discipline that came out of Google's SRE organization (the Site Reliability Engineers who run Google's production systems), and its whole idea fits in one sentence: a design isn't real until it has numbers on it. Not "we'll add a caching layer" — how many gigabytes of cache, holding what, at what hit rate? Not "we'll scale horizontally" — how many machines, and what resource on each machine runs out first?
The NALSD method: design the system for one machine first, then scale it up with arithmetic at every step. QPS (queries per second — request rate), bytes stored, IOPS (input/output operations per second — how many reads/writes a disk can physically do), machine counts. Here's the flavor: "100,000 reads per second, each one a random disk read. A spinning disk does about 200 random reads per second, so that's 500 disks doing nothing but serving reads — clearly wrong. An SSD does about 100,000 random reads per second, so a handful of SSDs cover it — or, better, the hot set is 40 GB, which fits in RAM on one big machine, replicated three times for redundancy." Four sentences, and the architecture chose itself. That's what "non-abstract" means.
Google did something rare for the interview world: it published the method. The Site Reliability Workbook — Google's own book on running production systems — contains a chapter introducing NALSD, and it walks through a worked example in exactly the style of the interview: designing a system that joins ad-click logs with search-query logs to compute click-through rates. The chapter starts with a design that fits on one machine, shows with arithmetic why one machine can't keep up — the log rate exceeds what one disk can write and seek through — and then scales up step by step: sharding the logs, computing how many disks the IOPS demand, how many machines the disks imply, and what the failure of any one of them does. Google's SREs use this method for real capacity planning, and the interview is a compressed version of the same exercise. If you want to know what the Google grader is trained to look for, the training material is sitting in public: every design claim is followed by the arithmetic that earns it.
The other Google signature: decisions tied to SLOs. An SLO — service level objective — is a promised target, like "99.9% of reads complete in under 100 ms." Google-calibrated candidates use the SLO as the arbiter of design arguments: "at 99.9% availability I can serve this from one region and eat the occasional blip; if you want 99.99%, that's under an hour of downtime a year, and now I need multi-zone redundancy and automated failover, which is why I'm asking which one we're promising." The SLO turns a taste argument into an engineering argument. Ask for the target early, and spend from it like a budget.
Palantir invented a round the rest of the industry doesn't have. The decomposition interview — "decomp" — hands you a vague operational problem, not a software prompt. Real examples candidates report: track the spread of an infection through a contact network. Assign airline gates at a hub airport. Plan the response to a flood. There is no clean answer, no expected architecture, and — this is the part that disorients people — often no explicit request to build software at all.
What's graded is everything that happens before architecture: requirements discovery (what actually matters here? what does "track" even mean?), stakeholder identification (who touches this system — nurses entering cases, epidemiologists querying clusters, an operations center making quarantine calls?), data modeling (what are the entities and relationships — people, contacts, test results, time — and how do they change?), and v1 prioritization (of everything you could build, what's the smallest version that helps someone tomorrow?).
A villager tells you: "People keep drowning crossing the river." A junior engineer hears "build a bridge" and starts computing beam loads. A senior engineer asks: who crosses, when, carrying what? Maybe the answer is a bridge. Maybe it's a ferry, a rope line, or moving the market so nobody needs to cross at dawn when the current is worst. The decomp round grades the questions, not the bridge — because Palantir's actual business is being dropped into a hospital or a factory with exactly this kind of problem statement. The analogy breaks in one way: in the interview you do eventually design something. But the design is graded on how well it fits the problem you uncovered, not on its architecture.
The decomp's reputation comes from a specific, repeatedly-told failure pattern: candidates with elite algorithmic backgrounds bombing it. The retellings rhyme. One candidate, given the infection-tracking prompt, recognized a graph problem immediately — contact network, nodes and edges — and spent forty minutes on an elegant traversal-and-clustering approach. Technically flawless. The interviewer's closing questions: who enters the contact data? (Nurses, by hand, mid-crisis — so the real bottleneck is a data-entry flow that works in ninety seconds.) How fresh is it? (Days stale — so real-time graph analytics computes precise answers to outdated questions.) What does v1 look like? (Silence.) The candidate had solved a problem nobody defined and prioritized nothing. Palantir keeps the round precisely because it's uncorrelated with algorithmic skill — it predicts something different: whether you can walk into a mess and find the problem worth solving. Candidates who pass tend to describe spending the first third of the interview asking questions and the interviewer visibly relaxing as they did.
Two variations to know. Most loops pair the decomp with a learning interview: you're dropped into something unfamiliar — typically a small codebase or a proprietary library, documentation included — and graded on how fast you orient yourself and build something real with it, because the job is permanent unfamiliar territory. The FDSE track (Forward Deployed Software Engineer — the ones embedded at customer sites) adds two more: a prototype exercise against genuinely messy data (think inconsistent, half-broken CSVs — the grading is on pragmatism, not polish), and a pitch to a deliberately skeptical "client," where the grade rides on whether you communicate honestly under pushback — admitting what your prototype can't do without collapsing.
Anthropic's loop, as candidates consistently describe it in 2024–2026: a recruiter conversation; a 90-minute practical coding screen (CodeSignal-style, in-browser); a hiring-manager deep dive on your background; then a roughly four-hour virtual onsite — two coding rounds (at least one usually touching concurrency: threads, locks, race conditions), one system design round, and one values round.
The screen deserves a warning label: it is explicitly not LeetCode. No dynamic-programming riddles. It's a practical build — implement a small working system against a spec that keeps extending — and it grades speed plus code that bends without breaking: can you ship a clean working version of part one fast enough to survive parts two, three, and four stacking requirements on top? Candidates who prepared with algorithm grinding report running out of time; candidates who write straightforward, extendable code report the opposite. The values round is also real, not a formality: Anthropic is an AI-safety company, and the round expects you to think out loud about AI risk and why this work matters — articulately and genuinely. Rehearsed slogans read worse than honest, considered uncertainty.
But the round to retool for is system design, because the flavor is infrastructure and LLM serving, not consumer products. Reported prompts: design the serving system behind a Claude-style chat product; design an LLM inference API; design search over a billion documents at high queries-per-second; design rate limiting for an LLM API. To play on this field you need a vocabulary most feed-and-fanout veterans lack:
A GPU serving chat requests is a restaurant where the constraint isn't the kitchen — it's the tables. Each conversation occupies a table (KV-cache memory) for its entire meal, and a long conversation is a camper who sits for three hours nursing one coffee. The kitchen (compute) could cook for a hundred diners, but if forty campers hold every table, new guests queue at the door — that queue is your TTFT going bad. Continuous batching is the maître d' seating new guests the moment any table frees up, instead of waiting for the whole dining room to empty. The analogy is honest about the core point — memory occupancy, not cooking speed, caps the restaurant — but it breaks at the edges: real schedulers can evict a camper mid-meal and re-seat them later (recomputing their state), which no restaurant survives.
Grading in the Anthropic design round leans on three axes. Cost — GPUs are staggeringly expensive, so "how many GPUs, and what does a million tokens cost us?" is a live question; a design that doubles utilization is worth real money and the interviewer knows the numbers. Latency percentiles — not "is it fast" but "what's the p99 TTFT, and what happens to it when a giant prompt lands in the batch?" (p99 = the time the slowest 1 in 100 requests sees — the tail your unluckiest users live in). Failure modes — a GPU dies mid-generation: is the conversation lost? A traffic spike arrives: who gets queued, who gets shed, does your rate limiter protect the p99 for everyone else?
| Company | Design round shape | Prompt style | Graded hardest | The trap |
|---|---|---|---|---|
| Meta E5 | One round; role tag picks System Design or Product Architecture | "Design a web crawler" / "Design Messenger" | Iteration under 100x pushes; one deep area; numbers and opinions from memory | Complete-but-shallow → E4 down-level |
| Google L5 | Candidate-driven; ~3/10/25/7 pacing; quiet interviewer | Deliberately vague slice: "Docs commenting" | The 25-min unprompted deep dive; NALSD math; SLO-tied choices | Reading silence as failure; hand-wavy boxes with no arithmetic |
| Palantir | Decomp round (+ learning interview; FDSE: messy-data prototype, skeptical-client pitch) | Vague operational problem, maybe not even software | Requirements discovery, stakeholders, data model, v1 prioritization | Jumping to algorithms or architecture before defining the problem |
| Anthropic | 90-min practical screen; onsite: 2 coding (concurrency), 1 design, 1 values | Infra/LLM serving: inference API, 1B-doc search, token rate limits | Cost, latency percentiles, failure modes; LLM-serving mechanics | Consumer-product prep; LeetCode grinding; slogan answers in the values round |
First, the Meta fork, because it's the decision people most often get wrong. The rule: pick the room where you have scars. If your production life has been pipelines, storage, queues, and infrastructure, pick System Design — you have real depth to deploy when the technical-excellence competency comes calling. If you've built product surfaces — APIs consumed by real clients, mobile or web apps, anything real-time — pick Product Architecture, because you can speak from experience about pagination contracts, reconnect logic, and what clients do when the network flakes. Do not pick Product Architecture because feeds feel more familiar than crawlers as a user. Familiarity as a user is worth nothing; the round grades the parts users never see.
Then point the drills at the exam:
Each company's senior probe has a signature sound. Learn to recognize which exam you're in from the question alone:
A junior friend says: "System design is system design — why would I prep differently for Meta versus Palantir?" Set them straight in five or six sentences.
The core skill transfers, but each company tests a different edge of it to failure. Meta gives you one design round that decides your level, pushes your design until it breaks with "now 100x users," and expects numbers and opinions from memory — so you drill speed and iteration. Google goes the other way: a quiet interviewer, a vague prompt, and more than half the time in one deep dive where every claim needs arithmetic — machines, disks, QPS — so you drill depth and math. Palantir barely asks about architecture at all: they hand you a messy real-world problem and grade whether you find out who needs what before building anything — a round that famously fails brilliant algorithm people. And Anthropic's design round is about serving language models — GPU memory, time-to-first-token, cost per million tokens — a domain with its own vocabulary you have to learn first. Same muscles, four different sports.
A friend has 6 years of experience: Kafka pipelines, a sharded MySQL fleet, an internal metrics system. She's also built a couple of GraphQL APIs "on the side" and once shipped a small React dashboard. She's applying to Meta E5 and asks: System Design or Product Architecture? Make the call and justify it in three sentences.
System Design, without hesitation. The round's technical-excellence competency demands at least one area of genuine E5 depth, and her scars — pipelines, sharding, metrics — all live on the backend exam; on that terrain she can deep-dive from experience, not from reading. Product Architecture would grade client-server contracts, live-update mechanics, and client failure handling — areas where two side projects give her breadth but no depth, and depth is exactly what decides E5 versus E4. Choose the exam where the hardest question lands on your strongest ground.
NALSD drill. A log-ingestion service receives 50,000 writes/sec, each 1 KB. It must also serve 20,000 random reads/sec of recent entries. You have machines with spinning disks (~200 random IOPS each, ~150 MB/s sequential) and 128 GB RAM. Do the arithmetic: what handles the writes, what handles the reads, and roughly how many machines before replication?
Writes: 50K × 1 KB = 50 MB/s — sequential appends, so a single disk's ~150 MB/s absorbs it comfortably; writes are not the problem. Reads: 20,000 random reads/sec ÷ 200 IOPS = 100 spinning disks doing nothing but seeking — clearly the wrong tool. So compute the hot set instead: if reads target the last hour, that's 50 MB/s × 3600 s = 180 GB — too big for one 128 GB machine, but two machines (or one, if reads skew to the last ~40 minutes) hold it in RAM, where 20K reads/sec is trivial. Answer: ~2 machines serving reads from memory, appending to disk for durability, ×3 for replication ≈ 6 machines. The senior signal is the pivot: the IOPS math didn't produce a disk count to buy — it proved disks were the wrong medium and redirected the design to RAM.
Decomp drill. Prompt: "A hub airport keeps having gate-assignment chaos when flights are delayed. Help." Spend 10 minutes producing only three artifacts: a stakeholder list, a rough data model, and a one-sentence v1. No architecture allowed.
One good shape. Stakeholders: gate agents (need tonight's assignments and changes, on their feet, on a phone), airline ops controllers (make the reassignment decisions), ground crews (fuel, baggage — need lead time before a gate change), pilots (need the gate before descent), passengers (downstream, via displays). Data model: flights (id, aircraft type, scheduled/estimated times), gates (id, size class, which aircraft fit, adjacency), assignments (flight→gate, time window, status), constraints (international flights need customs-adjacent gates; wide-bodies need large gates), delay events (flight, new estimate, timestamp). v1: a live board showing current assignments with conflicts highlighted the moment a delay lands — no optimizer, no auto-assignment, just making the collision visible ten minutes sooner than the phone calls do. The trap this avoids: leaping to "gate assignment is a scheduling optimization problem" and building a solver for controllers who first need to see the conflict — the exact algorithm-first mistake the round is designed to catch.
Anthropic drill. Your LLM API limits each org to 60 requests/minute. An org sends 60 requests of ~100K tokens each; another sends 60 of ~200 tokens. Explain concretely why this limiter fails, then sketch the replacement in 3–4 sentences with numbers.
The failure: both orgs are "at their limit," but the first is consuming ~6M tokens/minute of GPU time and the second ~12K — a 500x difference the limiter can't see. The heavy org can saturate your GPUs, wrecking p99 TTFT for everyone, while remaining fully compliant. Replacement: a token-per-minute budget per org — say 1M tokens/min — implemented as a token bucket refilling at ~16.7K tokens/sec; on each request, debit the known prompt tokens plus the request's max_tokens cap up front (the honest worst case, since completion length is unknown at admission time), then refund the unused portion when generation finishes. Requests that would overdraw the bucket queue or get a 429 with a retry-after. Now the limit measures what actually costs money and GPU time — tokens — and one org's giant prompts can't eat the tail latency of everyone else's.
It's your first on-call week at a new job. A ticket lands: "Your app takes four seconds to open. Unusable." The user is in Sydney. You check the dashboards: your servers, in Virginia, answer every request in under 50 milliseconds. You try the app yourself: instant. Nobody is lying. And yet almost four seconds are missing between that user's thumb and your perfectly healthy server.
Those missing seconds live in a part of the system most engineers never look at: the journey. Before a single line of your code runs, the user's phone has to find your server's address, open a connection, prove nobody is eavesdropping, and only then ask its question — and several of those steps are a full trip across the Pacific and back. Each trip has a price set by physics, and physics doesn't do discounts.
Interviewers love this territory. "What happens when you type a URL and press Enter?" sounds like a warm-up, but at the senior level it's a depth gauge. A junior answer names the steps. A senior answer prices them — where the round trips pile up, what gets cached where, and how that math explains half of modern architecture: CDNs, keep-alive, WebSockets, gRPC. Let's follow one request the whole way.
You typed hldgym.com. That's a name for humans. The internet routes traffic by IP address — a number like 203.0.113.7 that identifies a machine. The translation service is DNS — the Domain Name System — the internet's phone book, except there is no single book. There's a chain of people who each know a little.
First your browser checks its own memory: looked this up recently? Then the operating system checks its cache. If both are empty, your phone asks a recursive resolver — a server, usually run by your ISP or a public service like Google's 8.8.8.8 or Cloudflare's 1.1.1.1, whose whole job is doing lookups for you. "Recursive" means it chases the answer down, however many hops it takes.
If the resolver doesn't have the answer cached, it climbs a hierarchy. A root server — one of 13 named server clusters, mirrored worldwide — says, "I don't know hldgym.com, but here's who runs .com." A .com TLD server (TLD = top-level domain) says, "here's who runs hldgym.com's own name servers." Finally the authoritative name server — the machine family that actually holds the record — returns the IP address.
Scroll the diagram off the screen. Now, from memory: list every machine a cold lookup touches, in order, and what each one answers. The whiteboard will ask for this chain unprompted, so produce it unprompted.
Browser cache → OS cache → recursive resolver (the only machine your phone ever talks to) → root server ("ask .com") → .com TLD server ("ask hldgym.com's own name servers") → authoritative name server (the IP, plus a TTL) → back through the resolver, which caches it, to your phone. Warm cache: one hop, phone to resolver, ~1 ms.
Every answer travels with a TTL — time to live — a number of seconds meaning "you may cache me this long before asking again." With billions of lookups a day, DNS only survives because nearly all of them are cache hits somewhere. And TTL is a genuine trade-off. A long TTL (24 hours) buys fast lookups and resilience — even if your DNS provider has a bad hour, cached answers keep working — but when you change the record, the old answer haunts caches for up to a day. A short TTL (60 seconds) lets you move fast but drains your safety cushion: caches empty quickly, and every resolver on Earth comes knocking again.
The operational habit that falls out: before a planned migration, lower the TTL a day or two in advance, move, verify, then raise it back. Changing the TTL at migration time is too late — the old value, with the old long TTL, is already sitting in caches worldwide.
A recursive resolver is a hyper-organized receptionist. You ask, "where does Priya Sharma live?" She checks her notebook first — she wrote it down last week, with "valid until Friday" in the margin (the TTL). If the note has expired, she works the chain: the national directory ("try the Maharashtra office"), the state office ("try her housing society"), then the society office, which knows the flat number. You only ever talked to the receptionist. Where the analogy breaks: there are millions of receptionists like her, each keeping her own notebook — which is why the offices at the top of the chain don't collapse under the load.
On 21 October 2016, Twitter, Netflix, Reddit, Spotify and GitHub all "went down" for millions of people across the US East Coast — except none of them was down. Their servers were healthy. What broke was Dyn, the company running authoritative DNS for all of them. The Mirai botnet — roughly a hundred thousand hacked webcams and DVRs, conscripted through factory-default passwords — hammered Dyn's DNS servers in three waves through the day. A site whose name won't resolve might as well not exist. Users whose resolvers still held cached answers kept browsing — until the short TTLs expired minutes later and the caches drained. The lesson stuck industry-wide: big companies now run two independent authoritative DNS providers in parallel. When an interviewer asks "what are this design's single points of failure?", DNS is the one almost everyone forgets — because it fails so rarely that people forget it can.
One more piece of plumbing, because our next story needs it. DNS tells the world which address a name maps to. BGP — Border Gateway Protocol — is how the world knows which wires lead to that address: every network continuously announces "these IP blocks live here, route through me." DNS is the phone book; BGP is the road signs. And the road signs enable the internet's best magic trick, anycast: announce the same IP address from a hundred cities, and routing quietly delivers each user to the nearest copy. That is how the "13" root server addresses are really some two thousand anycast instances — many of them whole clusters — spread across the planet — the top of the DNS ladder never melts because it is everywhere at once — how 8.8.8.8 answers in a few milliseconds on every continent, and how serious DNS and CDN operators keep even the cold lookup a short trip. (Your phone may reach its resolver over DoH — DNS wrapped in HTTPS — these days; the chain behind it is unchanged.) Now watch a company take down both the phone book and the road signs at once.
On 4 October 2021, Facebook, Instagram and WhatsApp vanished for about six hours. During routine maintenance, a command meant to audit capacity on Facebook's backbone — the private network linking its data centers — instead disconnected all of them; the audit tool that should have blocked the command had a bug. Then the design ate itself: Facebook's DNS servers were built to withdraw their BGP announcements whenever they couldn't reach the data centers, on the theory that a cut-off DNS server shouldn't advertise itself as healthy. The backbone died, so every DNS server concluded it was unhealthy, and all of them pulled themselves off the internet at once. Nothing could resolve facebook.com — including Facebook: internal tools ran on the same domains, so engineers couldn't fix it remotely and had to drive to data centers, slowed by degraded badge systems. Meanwhile every phone on Earth kept retrying lookups — Cloudflare's 1.1.1.1 resolver saw roughly thirty times its normal query load for Facebook's domains. Two senior lessons, gift-wrapped: automation that removes capacity on failed health checks needs a bounded blast radius (withdrawing one unhealthy site is failover; withdrawing all of them is self-destruction), and you need an out-of-band way into your systems that doesn't depend on them being up.
Fine. Your phone holds an IP address. Note what it has not done yet: sent a single byte to your server.
Underneath everything, the internet moves packets — chunks of data, a kilobyte and change each — and its delivery promise is "best effort," a polite way of saying "no promise at all." Packets get dropped when routers are busy. They arrive out of order because they took different roads. For a cat video frame, who cares. For the bytes spelling "transfer ₹50,000," you care enormously.
TCP (Transmission Control Protocol) builds a trustworthy pipe out of this mess. Both sides number every byte, acknowledge what arrived, re-send what went missing, deliver data in exact order, and slow down when the network shows signs of choking (that's congestion control — collective politeness, enforced in code). But first, the two sides must introduce themselves via the three-way handshake: your phone says "can we talk? here's my starting sequence number" (SYN), the server replies "yes — here's mine, and I heard yours" (SYN-ACK), your phone confirms (ACK). One full round trip before any real data moves.
What's a round trip worth? Light in fiber covers about 200 km per millisecond, and the cable path from Sydney to Virginia is roughly 15,000 km — with real-world routing, a round trip (RTT) of about 200 ms. Nothing you deploy makes that number smaller. You can only make fewer trips, or shorter ones.
| Between | Typical round trip |
|---|---|
| Two servers, same data center | ~0.5 ms |
| Same city | 1–5 ms |
| Across a continent | 60–80 ms |
| Across the Atlantic | ~80 ms |
| India ↔ US East, Sydney ↔ Virginia | ~200–250 ms |
So why would anyone skip TCP's guarantees? Because the guarantees cost time, and some data expires faster than it can be repaired. UDP (User Datagram Protocol) is the alternative: no handshake, no ordering, no retransmission. Fire the packet and hope. Reckless — until you notice what kind of data suits it. A video-call frame from two seconds ago is worthless; re-sending it means freezing the call to deliver stale pixels. A game position update is obsolete the instant a newer one exists. A DNS query fits in one tiny packet — paying a round trip of handshake to send it is absurd, so DNS just sends, and re-asks on the rare miss. HTTP rides TCP because pages must arrive complete and correct. Live video, voice, games and DNS ride UDP because fresh beats perfect.
TCP is a phone call: you dial, hear "hello?", say "hi, it's me," and only then talk — and mid-conversation you can say "sorry, repeat that?" until you've heard everything, in order. UDP is dropping postcards into a mailbox: cheap, no setup, no confirmation, and if one goes missing you'll never know. Where the analogy breaks: postcards take days, which makes UDP sound slow — in reality both take milliseconds, and UDP is the faster one precisely because it skips the pleasantries. The difference is the promise, not the speed of the messenger.
Handshake done: you have a reliable line. One problem — every router, ISP, and coffee-shop Wi-Fi box along the path can read the whole conversation, and could quietly rewrite it.
TLS (Transport Layer Security — the S in HTTPS) fixes two separate problems; keep them separate in your head. First, encryption: the conversation is scrambled so nobody along the path can read it. Second — the part people forget — identity: how do you know the machine answering is actually hldgym, and not the coffee-shop router impersonating it? The server presents a certificate — "this public key belongs to hldgym.com" — digitally signed by a certificate authority, one of a small set of organizations your browser already trusts. Encryption without identity would just be a beautifully private conversation with an imposter.
Mechanics, one level deep: the two sides use slow asymmetric cryptography (public/private key math) to agree on a shared secret, then switch to fast symmetric encryption for the actual data. The negotiation costs round trips: TLS 1.2 needed two extra; TLS 1.3 got it to one, with session resumption for repeat visitors — even a "0-RTT" mode that sends the first request alongside the handshake. But 0-RTT carries a price the deleted trip used to pay: that first flight can be captured and replayed by someone on the path, so servers only accept requests in it that are safe to receive twice. This chapter opened on Bob Braden fighting to delete this very round trip in 1994 — T/TCP died because it dropped the proof the trip was buying. Every scheme that has deleted it since, from TCP Fast Open's cookie to 0-RTT's replay-safe-only rule, has had to buy that proof back some other way. The trip was never waste. It was evidence.
Now total the bill for a Sydney user's first request, at ~200 ms per round trip: DNS (cold) ≈ one trip to wherever the authoritative servers sit, TCP = one, TLS 1.3 = one, then the HTTP request itself = one more. Four round trips ≈ 800 ms before the first response byte — for a server whose own dashboards show a 50 ms response. A real page then fetches dozens of images and scripts on top. That is where the missing seconds go, and no server optimization will find them, because they were never on the server. One vocabulary note, because dashboards speak it: the median request (p50) rarely pays this bill — most visitors are warm somewhere. The Sydney first-timer lives out at p99, the latency your slowest 1% of requests exceed. Averages hide oceans; percentiles show them.
And the bill so far prices only the first byte. TCP does not trust a brand-new connection with full speed: it starts with about ten packets in flight — roughly 14 KB — and doubles the allowance each round trip as acknowledgements come back. This is slow start, and it means a 1 MB page costs five or six more round trips after every handshake is done. It is the second thing keep-alive buys: a warm connection has already grown its sending window. It is also why performance teams fight to fit the first meaningful paint into the first ~14 KB — those bytes arrive a full trip ahead of the rest. (How a sender should slow down again on loss — loss-based CUBIC versus model-based BBR — is a live argument; the cautious start is not.)
Diagram off the screen. Reconstruct the first 600 ms from memory: three round trips, what crosses the ocean in each, and the running clock. Say it out loud at interview pace before you write it.
Trip one, TCP: SYN out, SYN-ACK back — a reliable line at ~200 ms. Trip two, TLS 1.3: ClientHello out; certificate and keys back — identity plus encryption at ~400 ms. Trip three, HTTP: GET out, 200 OK back — first byte at ~600 ms. A cold DNS lookup paid a trip before all of it, and slow start rations the bytes after it.
This one diagram quietly explains a family of architecture decisions. Why do CDNs put servers in every city? To shrink the RTT that gets multiplied. Why do load balancers "terminate TLS at the edge"? To do the expensive handshakes near the user, then reuse warm connections to the backend. Why did Google invent QUIC? To collapse the number of trips. Different tools, one enemy.
Time to first byte = (setup trips) × (round-trip time). Physics prices the trip; software only picks the count. Run it forward and this chapter falls out. Delete trips: keep-alive, connection pools, TLS 1.3, session resumption, QUIC's merged handshake. Shorten trips: CDNs, edge TLS termination. Slow start stretches the rule past the first byte — the first ~100 KB is still trip-priced, which is why a warm connection wins twice. And know where it breaks: once the window is open, a big transfer is bandwidth-bound, and moving a 2 GB file closer barely helps. At the whiteboard, say the kernel first and derive the tool from it. Reciting tools without the kernel is the junior tell.
At last, the actual exchange. HTTP is plain, structured text: a request is a method plus a path, then headers (key-value metadata), a blank line, and an optional body. The response mirrors it, opening with a status code.
GET /feed?cursor=abc123 HTTP/1.1
Host: hldgym.com
Accept: application/json
Cookie: session=9f2c8a…
HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: max-age=60
{"items": ["…"]}
Status codes come in families, and the families matter more than the members: 2xx success; 3xx "look elsewhere"; 4xx "your request is wrong" — the client's fault; 5xx "I failed" — the server's fault. The split drives real behavior: a 4xx usually shouldn't be retried — the same wrong request fails the same way — while a 5xx might be a passing overload. But learn the two exceptions, because the senior questions live there. 429 Too Many Requests is a 4xx that exists precisely to be retried: later, more politely, honoring its Retry-After. And retrying a 5xx blindly can be worse than not retrying at all — a 500 on "transfer ₹50,000" does not tell you the transfer didn't happen; the write may have committed and only the response died, and a naive retry pays twice. A retry is only safe when the request is idempotent — built so that running it twice leaves the world exactly as if it ran once. That word gets a whole chapter later in this book; it decides who may retry what. The headers are quiet workhorses too: Cache-Control powers every caching layer you'll meet later in this book, Cookie and Authorization carry identity. When we say "the CDN respects your caching policy," we mean it reads these headers.
The original HTTP/1.0 did something that should now horrify you: it closed the TCP connection after every request. All those handshake round trips, paid again per image. HTTP/1.1's keep-alive fixed that — the connection stays open, the setup tax is paid once and amortized over many requests. Reuse matters even more between your own servers: backends keep pools of warm connections to each other and to databases, because at thousands of internal calls per second, re-handshaking would burn the whole latency budget. "Reuse warm connections" is one of the highest-value-per-word sentences you can say in an interview.
But HTTP/1.1 kept a painful rule: one request at a time per connection. A modern page needs dozens of resources, so everything queued behind everything else — head-of-line blocking: the guy at the front of the line holds up everyone behind him. Browsers coped with a hack: up to six parallel connections per site. Six lines instead of one.
HTTP/2 (2015) fixed it properly — at one layer. One connection carries many streams, chunks interleaved, so a big slow download no longer blocks a small urgent one (multiplexing, plus header compression). But a flaw hid one layer down. TCP promises to deliver bytes in order — it has no idea streams exist. When one packet is lost, TCP holds back everything after it until the retransmission arrives, including bytes of unrelated streams. On a lossy network — hello, mobile — HTTP/2 can be slower than HTTP/1.1, whose six separate connections at least failed independently. The queue vanished from HTTP and reincarnated inside TCP. Name the law once, because this book will keep meeting it: any promise of in-order delivery is a queue, and a queue has a head — whoever holds it blocks everyone behind. TCP promised order across everything on the connection; that promise is the jam.
How do you fix a queue that lives inside TCP? You can't patch TCP — it's baked into thirty years of operating systems and network boxes. The freezing even has a name: ossification. Middleboxes — NATs, firewalls, traffic "accelerators" — drop protocols they don't recognize and rewrite the TCP fields they do, so TCP can no longer evolve and nothing new can deploy beside it. SCTP, a genuinely better transport standardized back in 2000, never made it onto the open internet for exactly this reason — it lives on inside telecom cores, and tunneled over UDP inside WebRTC, but natively the middleboxes starved it. So you go around it. HTTP/3 runs on QUIC, a transport Google built on top of UDP — the only other thing every middlebox passes — using it as raw, promise-free material and implementing reliability, ordering and congestion control itself, per stream. Lose a packet from stream B, and only B waits; A and C keep flowing. QUIC also bakes TLS into its handshake — setup and encryption in one round trip, 0-RTT on resumption, carrying the same replay rule — and encrypts almost all of its own headers, so no middlebox can ever grow dependencies on them: armor against the next thirty years of ossification, not just privacy. It identifies connections by ID instead of by network addresses, so your download survives walking out of Wi-Fi range onto mobile data; and it lives in userspace, so fixes ship in browser updates instead of OS decades. The shape to remember: HTTP/2 removed the queue at the application layer; HTTP/3 removed it at the transport layer.
Everything so far shares one silent assumption: the client speaks first. But suppose your app is a chat app, and somewhere in Virginia a message has just arrived for your user. The server has news, and HTTP has no way for a server to call you. So… how does your phone ever find out?
This is one of the most reliably-asked topics in system design interviews, because every interesting product — chat, notifications, live scores, ride tracking — hits it. There are four standard answers, and each exists because the previous one hurt.
Short polling: the client asks on a timer. "Anything new?" every five seconds. Trivially simple, works through any network gear ever built — but news waits up to a full interval, and the waste is spectacular: a million users polling every 5 seconds is 200,000 requests per second, almost all answered "no."
Long polling: same question, patient answer. The server doesn't reply — it holds the request open until news arrives (or ~30 seconds pass), then responds; the client immediately asks again. Delivery becomes near-instant and the "no" traffic disappears. The cost: every user occupies one open request at all times, plus constant hang-up-and-re-ask churn. Its historical superpower is compatibility — it's just slow HTTP, so it slips through ancient proxies — which is why it survives today mostly as a fallback.
Server-Sent Events (SSE): stop pretending. The client makes one request, and the server sends a response that never ends — a stream of small events over one held-open connection. One-directional (server → client), built into browsers with automatic reconnection, and because it's ordinary HTTP, load balancers, proxies and CDNs handle it without ceremony. The quietly correct answer for feeds, notifications, dashboards, live scores — and it's how LLM chat interfaces stream their token-by-token replies.
WebSockets: the full upgrade. The connection starts as an HTTP request with an "upgrade me" header, then sheds HTTP and becomes a raw two-way pipe — full duplex: both sides send anything, anytime, with a few bytes of overhead per message. The tool for genuinely bidirectional, low-latency traffic: chat, multiplayer games, collaborative editors, trading. The price: you now run a stateful service. Each server holds thousands of long-lived connections — an idle one costs roughly 50 KB of kernel buffers and application state, so ~100K connections per untuned box is a sane planning number — so load balancers must route sensibly, deploys must drain connections gracefully, a momentary blip triggers a reconnect stampede, and heartbeats and re-authentication are yours to build. Stateless HTTP forgives operational sloppiness; WebSockets do not.
| Pattern | Direction | Latency | Server cost | Plays nicely with infra? | Reach for it when |
|---|---|---|---|---|---|
| Short polling | Client asks | Seconds (the interval) | Wasted requests at scale | Works everywhere | Infrequent checks: job status, cron dashboards |
| Long polling | Server → client, simulated | Sub-second | One held request per user + churn | Works almost everywhere | Fallback where WebSockets/SSE are blocked |
| SSE | Server → client only | Sub-second | One held connection per user | Yes — plain HTTP | Notifications, feeds, tickers, LLM token streams |
| WebSockets | Both, full duplex | Tens of ms | One stateful connection per user + ops burden | Needs LB and proxy support | Chat, games, collaborative editing, trading |
The framing that reads senior: choose by direction and latency, then say the operational cost out loud. One-way push? SSE — and mention it rides plain HTTP through every proxy and CDN. Genuinely two-way and interactive? WebSockets — and mention drain-on-deploy and reconnect storms in the same breath. A status page checked occasionally? "Polling every 30 seconds is honest and correct here" — resisting the gold-plating is itself a senior signal. The junior tell interviewers listen for: "WebSockets for everything," which volunteers to run the hardest, most stateful option for problems that never asked for it.
WhatsApp's product is a persistent connection. Every phone running the app holds one always-on TCP connection to a WhatsApp server — no polling, ever — so a message is pushed the instant it arrives, without burning battery on "anything new?" requests. The legendary part is the density. The early team ran a heavily tuned stack — Erlang on FreeBSD, built around the ejabberd messaging server — and in 2012 published benchmarks showing a single server holding over two million concurrent TCP connections. Erlang made that sane: millions of isolated, lightweight processes at a few kilobytes each, so "one process per connected phone" is a design, not a fantasy. The efficiency shaped the company — when Facebook paid $19 billion for WhatsApp in 2014, roughly 450 million users were served by about fifty employees, around thirty of them engineers. The senior lesson: in persistent-connection systems, concurrent connections — not requests per second — are the scaling axis. When you say "WebSockets" in an interview, beat the interviewer to the follow-up: how many connections per box, what they cost in memory, and what happens when a deploy or a blip disconnects a million phones at once.
The story hides one more senior probe, so take it with you: a held connection only solves the last hop. Your friend's "send" lands on any stateless API server — but your socket lives on connection server 47 of 400. How does the message find it? Every persistent-connection system needs a routing tier: a session registry (which server holds this user's socket) or a pub/sub bus that every connection server subscribes to. WhatsApp's Erlang registry — one process per phone, addressable by name — is that answer; "one socket per phone" was never the whole design. When you say "WebSockets" at a whiteboard, the next question is "and how does a message find the socket?" — now you have the shape. The full design is a Part 3 chapter.
And connect the stampede to this chapter's kernel: every held connection is a handshake you paid once and have been amortizing ever since. A 30-second blip presents the whole invoice at once — millions of phones re-doing TCP, TLS and authentication in the same minute, against servers sized for the cheap steady state. Steady state is easy; the transitions are the design.
One boundary left. Everything above was the public internet: browsers, oceans, 200 ms round trips. Behind your load balancer is a different world — your services talking to each other at 0.5 ms RTT, but at brutal volume, because one user request commonly fans out into dozens of internal calls.
At the public edge, REST — resource-shaped URLs, HTTP verbs, JSON bodies — earns its throne. Every language, tool and intern speaks it; you can debug it by reading it; HTTP caching understands it. But between internal services, JSON's charms invert: it's text, so it's bulky; parsing burns CPU; and nothing checks that the user object one service sends still matches what the next expects — you find out at runtime, in production.
gRPC is the internal-traffic specialist. You define every call and message in a schema file (Protocol Buffers, or protobuf), and code generation produces typed client and server stubs in each language — calling another service looks like an ordinary typed function call, and misuse fails at build time — within your own binary. Across services the guarantee is different, and knowing the difference is what reads senior: two services are never compiled together, because every rolling deploy has version N and N+1 talking for hours. Protobuf's real cross-service contract is wire-level evolution discipline — fields are identified by number, old readers skip fields they don't know, and you may add fields freely but never renumber one, never change its type, never reuse a deleted number. What pages you at 2 am with protobuf is a reused field number silently decoding as garbage — quieter than JSON's loud parse error, which is why .proto files get code-review scrutiny like public APIs. On the wire, protobuf is compact binary — typically a few times smaller than the same JSON, and much cheaper to parse. Underneath, gRPC rides HTTP/2: warm, multiplexed, long-lived connections. It streams in both directions and has deadlines built in — each call carries its own time budget, so a slow downstream fails fast instead of piling up waiting callers. At tens of thousands of internal calls per second, the saved serialization and connection overhead is real CPU and real money.
The costs, honestly: browsers can't speak native gRPC (you need a proxy layer like gRPC-Web, which is why it stays behind the edge), binary payloads can't be debugged by squinting at curl output, and schemas demand discipline. So the pattern interviewers expect: REST at the public edge; gRPC for chatty internal paths — with the judgment call that plenty of successful companies run JSON internally forever, because gRPC's payoff arrives with call volume and team scale, not on day one.
This material surfaces constantly at senior loops — rarely as its own question, always as a depth probe inside a bigger design. A passing answer prices the network instead of hand-waving at it: it counts round trips, names what's cached where, and picks real-time transports by direction and latency with the operational cost said out loud. But the level is decided on the follow-ups, so practice the ladder, not just the first rung:
"Which of those steps disappear on the second request — and who remembered what?" Browser DNS cache, resolver cache, the kept-alive connection, the TLS session ticket, the already-grown congestion window. "And which single change removes the most trips for a first-time visitor?" Edge termination — the handshakes happen 10 ms away instead of 200.
"You shipped the edge. p50 is fixed, but mobile p99 is still 1.8 s — now what?" Packet loss on cellular plus TCP head-of-line blocking under HTTP/2; the direction of the fix is HTTP/3. "How do you prove that before migrating two production apps?" Measure from the client, split by geography and network type — the missing time was never on the server, so the server's dashboards cannot see it.
"10 million concurrent — what does the fleet look like?" ~50 KB per idle connection, ~100K per box, so ~100 servers — and capacity planned for the reconnect storm, not the steady state. "A message arrives for a user — how does it find their socket?" A session registry or a pub/sub tier across the connection servers. Raising the stampede and the routing tier before they ask is the senior signal.
"What actually breaks when you evolve a .proto?" Nothing at compile time across services — wire compatibility is the contract: numbered fields, additive changes, never reuse a number. "Your three-person startup's CTO wants gRPC everywhere, including the public API — talk them down." Browsers need a proxy layer, debuggability drops, and the payoff arrives with a call volume you don't have. That restraint is itself the senior answer.
Explain to a junior engineer, in four or five sentences, why the very first HTTPS request to a website is slow but the second one is fast.
Before the first request can even be sent, your device does three rounds of setup: it looks up the site's IP address with DNS, does the TCP handshake to establish a reliable connection, and does the TLS handshake to encrypt it and verify the site's identity. Each of those is at least one full round trip to a server that might be thousands of kilometers away, so at 200 ms per trip you've spent 600–800 ms before the server sees your request at all. The second request skips almost everything: the DNS answer is cached on your device, and the TCP+TLS connection is kept alive and reused, so your request rides an already-open, already-encrypted line whose sending window has already grown. That's why setup cost, not server speed, usually dominates first-load latency — and why so much of system design (CDNs, keep-alive, QUIC) is really about paying that cost less often or closer to the user.
You chose SSE for the live score strip. The interviewer leans in: "Why not WebSockets — one technology for every real-time need?" Defend the choice out loud first, at interview pace, in under 30 seconds. Then write the skeleton of what you said, including one cost you accept.
The data flows one way — server to millions of viewers — so I won't pay for a duplex pipe I'd use in one direction. SSE is plain HTTP, so every proxy, load balancer and CDN between us already handles it, and the browser reconnects on its own. WebSockets would make this a stateful tier: drain-on-deploy, reconnect storms, heartbeats — operational cost with no matching requirement. The cost I accept: if the product later needs client-to-server interactivity, I add a WebSocket path for that feature. A concession with a boundary beats loyalty to a technology.
A user in Sydney opens your site for the first time (nothing cached anywhere). Your only servers are in Virginia; RTT is 230 ms; TLS 1.3. Estimate time-to-first-byte, then name the two changes that cut it most.
List the four things that must happen before the first response byte. Each costs about one round trip.
Count trips: DNS ≈ 1 RTT, TCP handshake = 1, TLS 1.3 = 1, request/response = 1. About 4 × 230 ms ≈ 900 ms before the first byte. Biggest levers: (1) move the connection endpoint closer — a CDN or edge PoP (point of presence: a small server site near users) near Sydney turns the 230 ms RTT into ~10–20 ms for DNS, TCP and TLS, saving ~600 ms even if the origin fetch still crosses the ocean once — and serious DNS providers already answer via anycast, so even the cold lookup is a short trip; (2) make fewer trips — keep-alive plus TLS session resumption (or HTTP/3's combined handshake) collapses repeat-visit setup to zero or one trip. Both levers attack the network math; nothing about the server changed. One honesty check: ~900 ms buys the first byte, not the page — a 1 MB body pays roughly five more trips to slow start, which the same two levers also shrink.
Pick a transport — short polling, long polling, SSE, or WebSockets — for each, with a one-sentence justification: (a) a rider's map showing the driver's live location; (b) a collaborative document editor; (c) a page showing progress of a 20-minute report-generation job; (d) a live cricket score strip with millions of viewers.
For each: which direction does the data flow, and how stale can it be before it is useless?
(a) One-way, frequent, small updates: SSE — don't pay for a duplex pipe you use one way. (b) WebSockets — edits flow both directions with low latency, the exact case that justifies stateful connections. (c) Short polling every 15–30 seconds — the job is slow and staleness is harmless; anything fancier is over-engineering, and saying so is the senior move. (d) SSE — one-way fanout of tiny events over plain HTTP, which your proxies and CDNs can help distribute; a WebSocket per casual visitor would be millions of stateful connections carrying strictly one-way data.
Your API domain's DNS TTL is 30 seconds "so we can fail over fast." Your DNS provider suffers a one-hour outage. What happens, and how would you re-design? Compare with TTL 24 hours.
What is left in the world's resolver caches 60 seconds after the outage starts?
Within about a minute, every resolver cache holding your records expires and can't refresh — your domain effectively stops resolving for nearly everyone, and your healthy servers serve nobody. At TTL 24 hours, most active users' resolvers would have coasted through the hour on cache; only cold lookups would fail. But the fix isn't "crank TTL to a day" — that makes every migration and failover glacial. The senior design: two independent authoritative DNS providers serving identical records (the post-Dyn standard), each answering via anycast from many sites, a moderate TTL (5–15 minutes), and external monitoring that resolves your domain from outside your own network so you learn about failures before Twitter does.
Your chat product holds 5 million concurrent WebSocket connections. Estimate the fleet (use the chapter's planning numbers: ~50 KB per idle connection, ~100K connections per server), then describe the 60 seconds after a 30-second network blip disconnects everyone — and what you'd build to survive it.
Multiply first. Then ask what all 5 million clients do in the same second after the blip ends.
At ~50 KB per connection (kernel buffers + app state), 5M × 50 KB ≈ 250 GB of connection memory fleet-wide; at a comfortable 100K connections per server, ~50 servers (WhatsApp ran 20× denser, after years of tuning). The blip is the real exam: 5 million clients reconnect at once, and each reconnect is a fresh TCP + TLS handshake plus authentication — far costlier than the idle connection was — so the stampede can down servers that held steady state easily. Survival kit: client-side reconnect backoff with jitter (spread retries randomly over tens of seconds instead of one synchronized wave), cheap session-resumption tokens so reconnect auth is light, connection-count-aware load balancing, and capacity planned for the storm, not the steady state. Steady state is easy; the transitions are the design.
Your Sydney user's feed is fully personalized — a CDN can cache none of it. The team concludes "so the edge can't help us." Check that with numbers: edge PoP 15 ms from the user, a warm pooled connection from the edge to Virginia, origin compute 50 ms. Compare against exercise 1's ~900 ms.
Which of the four setup trips can happen at the edge? And what does the edge already hold open to Virginia?
The handshakes move to the edge: DNS, TCP and TLS at ~15 ms each ≈ 45 ms. The request then crosses the ocean once, on a connection the edge keeps warm — no handshake, window already grown: ~230 ms, plus 50 ms of origin compute. Total ≈ 325 ms against ~900 ms — roughly a 3× win at a 0% cache hit rate, because setup trips dominated the bill and setup is exactly what the edge absorbs. Caching is a second, separate win (the CDN chapter, later in Part 1). Now run the kernel somewhere it has never been: a geostationary-satellite user has ~600 ms RTT to everywhere, so the kernel predicts the web feels broken there regardless of bandwidth — and that low-orbit constellations were a latency project before they were a bandwidth one. It also predicts where the edge stops helping: the 2 GB download, where the window is open and bandwidth rules.
The interview's output mode is speech under pushback, so the exit test is spoken. You are done with this chapter when you can, without looking:
A user on a shaky metro connection taps "Pay ₹4,999" in your app. The spinner spins. Ten seconds. Timeout. Here's the ugly part: you don't know what happened. Maybe the request never reached your server. Maybe it arrived, charged the card, and the response died on the way back. The app can't tell the difference. Neither can the user — so they do what every human does with a frozen button. They tap it again.
Now you may have charged them twice. Two bank SMS alerts, one angry screenshot, one support ticket. Meanwhile, a partner's integration broke overnight because someone on your team renamed a JSON field "nobody used." And your feed is quietly showing some users the same post twice on page two — but only when the feed is busy, so nobody can reproduce it.
Three unrelated bugs? No. One bug, in three costumes. In each case an API behaved like a casual conversation when it needed to behave like a signed contract. An API is a promise you make to code you will never see, written by people you will never meet. Once someone integrates against it, you honor it for years — or you break strangers' software every time you "clean things up."
This chapter is about writing promises you can keep. Interviewers probe this hard — Meta has an interview track that grades it explicitly — because API design is judgment made visible.
REST's core idea fits in one sentence: URLs name things (resources — nouns), and HTTP methods say what you're doing to them (verbs). GET /orders/42 reads order 42. POST /orders creates one. DELETE /orders/42 removes it. The URL never contains the verb — /getUserOrders and /orders/42/delete smuggle the action into the noun and throw away everything HTTP already gives you: caching for reads, safe retries for some verbs, uniform tooling.
The table below is worth memorizing — especially the last two columns. Safe means the call changes nothing on the server. Idempotent means calling it five times has the same effect as calling it once. These decide which requests a client, proxy, or load balancer may automatically retry after a timeout.
| Method | Meaning | Safe? | Idempotent? |
|---|---|---|---|
GET | Read a resource | Yes | Yes |
PUT | Replace the resource at this exact URL | No | Yes — replacing twice with the same thing is one replace |
DELETE | Remove it | No | Yes — deleting twice still leaves it deleted |
PATCH | Partially update | No | Not guaranteed |
POST | Create / trigger an action | No | No |
Look at that last row. POST — the verb behind "create order," "send message," "charge card" — is the only one where a blind retry causes damage. Hold that thought; it comes back with money attached.
An API is a wall socket. Every appliance in the country is molded to that exact plug shape by manufacturers who never met the electrician — useful precisely because it's boring, standardized, never surprising. Change the socket shape and you don't inconvenience people; you brick every device in the nation, which is why countries that switched plug standards kept old sockets working for decades. Where the analogy breaks: sockets stay frozen forever, but your API must grow. Evolving the socket without breaking old plugs is what versioning solves — we'll get there.
A status code's first digit is a contract about whose fault it was and what to do next. 2xx: it worked. 4xx: the request is wrong — fix it before retrying, because the same request earns the same rejection. 5xx: the server failed — the request may be fine, and a retry with backoff is reasonable. Retry logic, monitoring dashboards, load balancers, and CDNs all branch on that first digit without reading your body. That's the point: machines react correctly without understanding you.
Within the classes, a dozen codes carry the meaning: 400 malformed request, 401 we don't know who you are, 403 we know exactly who you are and the answer is no, 404 no such resource, 409 conflict with current state, 422 well-formed but semantically invalid ("quantity: -3"), 429 slow down, 500 we crashed, 503 temporarily unable — try again shortly, 504 a server behind us timed out. Precision matters: 409 vs 500 is the difference between "don't retry, re-read the resource" and "retry with backoff."
Which brings us to the cardinal sin: returning 200 OK with {"error": "not found"} in the body. It feels harmless. It is sabotage. Monitoring reports a 0% error rate while users see failures. A CDN caches the "successful" error page. Client retry logic sees success and moves on. Every machine between you and the caller reads the status code, and you just lied to all of them.
Nobody returns a million rows in one response, so every list endpoint pages. The obvious design is offset pagination: GET /feed?page=3&per_page=20, translated to SQL as LIMIT 20 OFFSET 40. It's everywhere, and it has two failure modes that both get worse with success.
Failure one: it lies when data moves. Offset means "skip N rows from the top, counted right now." But a feed's top is a moving target. Between page one and page two, a new post lands on top — every row shifts down — and the last row of page one becomes the first row of page two. The user sees it twice. A deletion shifts everything up instead, and a row silently vanishes between pages. On a busy feed this isn't rare; it's constant.
Failure two: it gets slower the deeper you go. OFFSET 100000 doesn't teleport to row 100,001. The database walks 100,000 rows in index order and throws them away — every request. Cost grows linearly with depth: page 5,000 costs 5,000 times page one. A crawler walking your whole collection becomes a database load test you didn't schedule.
The fix is cursor pagination (keyset pagination). Instead of "give me page 3," the client says "give me 20 items after this exact position" — the position being a cursor the server returned with the previous page: an opaque token, typically base64, wrapping the sort values of the last item, like (created_at, id). The server turns it into WHERE (created_at, id) < (:ts, :id) ORDER BY created_at DESC, id DESC LIMIT 20. That's an index seek: jump straight to the position, read 20 rows. Constant cost at any depth — and insertions above your position can't shift anything, because your bookmark is attached to an item, not a row count.
Offset pagination is finding your place in a loose-leaf binder by page number while someone keeps inserting pages at the front — "I was on page 12" is worthless once everything renumbers. A cursor is a physical bookmark clipped to a specific page: insert all you want in front of it, it still marks exactly where you stopped. The catch is real, though — a bookmark can't "jump to page 47." Cursors give up random access, which a feed never needed but a paged admin table does.
Two details make cursors production-grade. First, the sort key must be unique and stable — created_at alone has ties, and a cursor pointing into a tie skips or repeats rows; the fix is a compound key with a unique tie-breaker: (created_at, id). Second, keep the cursor opaque. If clients can read ?after_id=812, they'll construct their own, and your internal sort encoding becomes a public contract you can never change. Twitter's timeline API is the classic example: it paginates by tweet ID (max_id/since_id) because a timeline's top moves every second — and because Twitter's Snowflake IDs embed their creation time (newer tweet = bigger ID), the ID itself is a perfect cursor.
Slack's Web API originally paginated the obvious way: numbered pages backed by MySQL LIMIT/OFFSET. It worked — until Enterprise Grid brought customers with hundreds of thousands of users in one workspace. A deep page of a huge member list meant MySQL scanning and discarding enormous numbers of rows per request, and membership churn between requests drifted the page boundaries, duplicating or dropping users. Slack's engineers documented the move to cursor-based pagination on methods like conversations.members: responses return an opaque base64 next_cursor encoding the last item seen, and the next request resumes from an index seek at that position. Bounded cost at any depth, stable pages under churn — and an opaque cursor kept them free to change what's inside it. Their guidance to consumers became blunt: if an endpoint offers cursors, use them.
Filters and sorts ride in query parameters: GET /orders?status=shipped&sort=-created_at (minus prefix = descending, a common convention). Two rules. Whitelist what's sortable and filterable — every sortable field promises an index; accept sort=anything and you've handed strangers a full-table-scan button on your primary database. Bind cursors to their sort — a cursor encodes a position within one specific ordering, so if the client changes sort or filters mid-walk, reject the stale cursor cleanly instead of returning garbage.
Now the case where a retry costs money. The root problem: a timeout is not an answer — it's the absence of one. After a timed-out POST, exactly one of three things is true: the request never arrived; it arrived and is still executing; it finished and the response was lost. The client cannot tell these apart from outside. Not retrying risks a lost order. Retrying risks a duplicate. GET, PUT, and DELETE escape the dilemma — idempotent, retry freely. POST is the problem child, and "charge this card" is a POST.
The fix: make POST idempotent by hand. The client generates a unique idempotency key — a random UUID — for each logical operation (this checkout, this transfer) and sends it as a header: Idempotency-Key: 8e03978e-40d5-.... Crucially, retries reuse the same key; a fresh key per attempt would defeat the whole mechanism. Server side: if the key is new, record it, execute, and store the response (status and body) against the key. If the key has been seen and the operation finished, skip execution and replay the stored response, byte for byte. Now a timed-out payment can be retried with zero fear: either the first attempt never landed and this one executes, or it landed and the client gets the original receipt back. Both paths end in exactly one charge.
The edge cases are where interviews (and production) are won. Retry races the original mid-execution: lock on the key; the duplicate gets a conflict error ("in flight, retry shortly"), never a second execution. Same key, different body: a client bug — reject loudly rather than replay a response to a different question. Storage: the key record and the operation's effect must commit atomically — same database transaction — or a crash between "charged the card" and "stored the key" reopens the exact hole you were closing. Keys get a TTL so the store doesn't grow forever; a day comfortably outlasts any sane retry window. Chapter 2.2 goes deeper into exactly-once and ledgers; this header pattern is the API-facing half.
It's a wire transfer with a reference number. "Transfer ₹50,000, reference RENT-MARCH" — and the phone line dies before you hear a confirmation. You call back and repeat the reference. The teller doesn't blindly send again; they check first: "That went through at 2:14, here's your confirmation." Sent once, however many times the line drops, because the reference names the operation, not the phone call — which is exactly why retries must reuse the key. A new reference per call would look like a new transfer.
Stripe moves payments over the public internet, where timeouts aren't an edge case — they're weather. So idempotency keys have been first-class in their API for over a decade, and their design is the reference implementation everyone copies. Clients send an Idempotency-Key header (docs recommend a random V4 UUID) on POSTs. Stripe saves the status code and body of the first request made with a key — whether it succeeded or failed — and replays that saved result for every later request with the same key, so a client retrying into an error sees the same error, not a confusing second outcome. Keys are pruned after 24 hours. Reuse a key with different parameters and Stripe rejects it rather than guess which request you meant; a retry racing the original mid-flight is refused with a conflict. Their engineering writing pushed the pattern into the wider industry — it's now an IETF draft header — and it's why "how do clients safely retry payments?" has a canonical answer you're expected to know cold.
Everyone frames versioning as a syntax debate — version in the URL (/v2/orders) or in a header (Api-Version: 2024-06-01)? URL versions are visible, trivially routable, cache-friendly; header versions keep URLs stable and allow per-request pinning. Fine. But that's the small question. The big one: when do you get to turn v1 off? The honest answer is approximately never. Somewhere, a cron job written in 2019 by a contractor who left in 2020 is calling your v1 endpoint, and it will keep calling until the company running it dies. Announce a shutdown date and you'll discover how many of your biggest customers are that cron job.
Accept "v1 is forever" and strategy inverts: the cheapest version to support is the one you never had to create, so the real game is evolving without breaking. Additive changes are safe: new endpoints, new optional request fields, new response fields — provided your docs order clients to ignore unknown fields. Breaking changes — renaming or removing a field, changing a type or a meaning — are the expensive ones. Most "we need v2" moments are actually "we want to rename things" moments; the senior move is to add the new field beside the old, deprecate the old, and swallow the aesthetic discomfort. You sweat v1's field names because they're permanent.
Stripe took "v1 forever" to its logical extreme. Versions are dates — the day a breaking change shipped — and every account is pinned to whatever version was current at its first API call. A 2016 integration still sees 2016-shaped responses today, no action required; you upgrade deliberately, and can test newer versions per-request via a Stripe-Version header. The engineering that makes decades of versions affordable: each breaking change is a small, self-contained version change module declaring how to transform requests and responses across that one boundary. Core code targets only the current version; a pinned-2016 request is walked forward through each module on the way in, and the response walked backward through the same chain into 2016 shape on the way out. Supporting an old version costs a stack of tiny transformations, not a fork of the codebase — which is why Stripe can afford a promise most companies can't.
Every public API eventually meets a client in a hot retry loop — buggy, greedy, or malicious, the traffic looks identical. The contract for refusing is 429 Too Many Requests, plus headers that let well-behaved clients avoid the wall entirely: X-RateLimit-Limit (your quota), X-RateLimit-Remaining, X-RateLimit-Reset (when it refills), and on the 429 itself, Retry-After — the trio GitHub's API popularized. Clients hold up their end with exponential backoff plus jitter: randomness in retry delays, so a thousand rejected clients don't all return in the same instant like a synchronized battering ram. Never answer a 429 with an instant retry. The algorithms behind the limits get their own chapter (1.12).
The status code says what class of failure; the body must say which failure and what to do. A good error body is boringly consistent across every endpoint:
{
"error": {
"code": "card_declined",
"message": "Your card was declined. Try another payment method.",
"request_id": "req_7f3aa81b2c",
"param": "payment_method"
}
}
Each field has a job. code is machine-readable and stable — clients write if code == "card_declined" branches against it, making it part of your contract; renaming a code is a breaking change. message is for humans and may be reworded freely — clients must never parse it. request_id is the underrated hero: it appears in the response and your logs, so "it failed, here's the ID" turns support archaeology into one log query. RFC 9457 ("problem details") standardizes a similar shape if you'd rather adopt a spec. And never leak stack traces — a security disclosure wearing a debugging costume.
So far the client asks and you answer. But "tell me when the payment settles" shouldn't mean polling GET /payments/42 every five seconds forever. A webhook flips the arrow: the consumer registers a URL, and you POST events to it — payment.succeeded, order.shipped — as they happen. Now you're the client, calling servers you don't control over the same unreliable internet, and the receiving side is where integrations quietly rot.
Providers deliver at-least-once: no 2xx back means retry, on a backoff schedule — Stripe keeps retrying for days. That one fact dictates all four receiver duties. One: verify the signature. A webhook URL is a public door; anyone can POST fake "payment succeeded" events at it. Providers sign each payload with a shared secret (an HMAC in a header) — check it before believing anything. Two: acknowledge fast. Return 200 the moment the event is durably queued; process afterward. Heavy work inline means timeouts, which mean retries, which repeat the heavy work — a retry storm you built yourself. Three: dedupe. Every event carries an ID (evt_...); keep a table of processed IDs and skip repeats — idempotency keys again, arriving from the other direction. Four: never trust arrival order. Retries and parallel delivery mean order.shipped can land before order.created. The robust pattern: treat the event as a doorbell, not a database — on receipt, fetch the resource's current state from the API instead of replaying event payloads into your tables.
Everything above assumed strangers calling over the public internet, where JSON-over-HTTP earns its keep: human-readable, curl-able, callable from any browser. But when service A calls service B inside your own walls ten thousand times a second, those virtues invert. Nobody reads those payloads — but every one pays JSON's tax: verbose text on the wire, real CPU burned parsing it. At internal volumes, that tax lands on your cloud bill.
gRPC — Google's open-sourced descendant of Stubby, the RPC framework behind their internal fleet — is the standard answer. Contracts live in .proto files: declare messages and methods once, and a compiler generates typed client and server code in every language your company speaks, so a mistyped field name becomes a compile error instead of a 3am incident. On the wire, protobuf encodes messages in compact binary — several times smaller than JSON, far cheaper to parse. Underneath, HTTP/2 multiplexes many concurrent calls over one long-lived connection and streams in both directions. Deadline propagation deserves special mention: a request's remaining time budget travels down the call chain, so when the user has already given up, the fifth service downstream stops working on a corpse. The honest trade-off: you lose curl-ability and easy browser access (grpc-web needs a proxy). Hence the boring, correct rule: REST at the edge, gRPC inside.
And GraphQL? One honest paragraph. Facebook built it because mobile screens needed deeply nested, screen-specific data — a post with author, first three comments, each with likes — and REST forced either N round trips or bloated one-size-fits-all responses. GraphQL lets each client ask for exactly the fields it needs from one endpoint — genuinely valuable when you have many diverse clients with fast-changing needs and a platform team to own the schema. The bill: HTTP caching mostly stops working (every query is a POST), rate limiting must become query-cost analysis because one nested query can be a thousand database hits, and naive resolvers breed N+1 query storms (fetch a list in one query, then fire one more query per item). For two internal services with stable needs it's over-engineering — say so in the interview, and say what would change your mind.
Meta's Product Architecture track grades API design as an explicit rubric line, and the grading principle is: client needs drive the shape. Strong candidates start from the screens — "the feed view needs post, author, like count, and a page of comments" — and derive endpoints, fields, and pagination from that, rather than decorating a database schema with URLs. Then they volunteer the unhappy paths — timeouts, retries, evolution — without being dragged there. Expect probes like:
Explain to a junior engineer, in four or five sentences, why a timed-out payment request is dangerous to retry — and how an idempotency key makes the retry safe.
A timeout doesn't mean the request failed — it means you never heard the answer, so the charge may or may not have happened. Retry blindly and you risk charging the customer twice; don't retry and you risk losing the order. So the app generates one random key for the whole operation and sends it as a header on every attempt, including retries. The server remembers each key it has completed, along with the response it sent: a new key executes normally, while a repeated key skips execution and replays the saved response. Either way the charge happens exactly once, so the client can retry as often as the network demands, without fear.
Design the pagination for a comments API on a busy post (thousands of new comments per hour), newest first. Specify the request, the response shape, the underlying query, and one sentence on why not offset.
Request: GET /posts/42/comments?limit=20&cursor=eyJ0cyI6... (cursor omitted on the first page). Response: the items, plus an opaque base64 next_cursor encoding the last item's (created_at, id), plus has_more. Query: WHERE post_id = 42 AND (created_at, id) < (:ts, :id) ORDER BY created_at DESC, id DESC LIMIT 20, backed by an index on (post_id, created_at, id) — the id tie-breaker prevents duplicates and skips when comments share a timestamp. Why not offset: constant insertions shift row positions, so offset pages repeat and skip comments — and deep offsets scan-and-discard linearly while the cursor's index seek stays constant-cost.
Code review this API sketch and list every sin with its fix: (a) GET /getUser?id=5 returns 200 with body {"error": "user not found"}; (b) POST /users/5/delete; (c) the activity feed uses ?page=3; (d) the versioning plan is "we'll just update the API and email the partners."
(a) Two sins: a verb in the URL, and a lie in the status code — make it GET /users/5 returning 404 with a structured error body; the 200-wrapped error blinds monitoring, poisons caches, and fools retry logic. (b) DELETE /users/5 — which also makes the operation officially idempotent, so clients may safely retry it after a timeout. (c) An activity feed is high-churn, so page numbers duplicate and skip items as rows shift — use cursor pagination over (created_at, id). (d) "Update and email" is a breaking change with extra steps; evolve additively (add fields, never rename or remove), tell clients to ignore unknown fields, and reserve explicit version pins for the rare true break — with the old version kept alive effectively forever.
A client POSTs a transfer with idempotency key k1 and times out. Enumerate the three possible server-side states at that moment, then give the server's correct behavior for a k1 retry in each — plus the fourth case: k1 arrives with a different request body.
State 1 — never arrived: no record of k1, so the retry executes as a first attempt. State 2 — still executing: k1 is recorded with no stored response; return a conflict ("in flight, retry shortly") — never start a second execution. State 3 — completed, response lost: replay the stored status and body verbatim, executing nothing. Fourth case: same key, different body is a client bug — reject with an explicit error, never replay, because the stored response answers a different question. Bonus point worth saying aloud: the key record must commit in the same transaction as the transfer, or a crash between them reopens the double-execution hole.
Your public API returns "amount": 49.99 — rupees as a float — and floating-point rounding is causing real accounting bugs. You want integer paise. Thousands of integrations depend on the current field. Ship the fix without breaking anyone.
Additive migration: add "amount_paise": 4999 beside the old field and keep both indefinitely. Deprecate amount in docs and SDKs, nudge laggards from the dashboard — but never remove it on a timeline you merely announced; the 2019 cron jobs don't read email. With date-pinned versioning, Stripe-style, you can additionally drop the float field from new-version responses via a version-change transform, so new integrations start clean while pinned ones stay whole. The junior answer is "rename it in v2, sunset v1 in six months"; the senior answer knows v1's sunset date is approximately never, so it makes the old contract harmless instead of trying to kill it.
Sketch a webhook receiver for a payment provider sending ~200 events/sec, with occasional retry storms when your endpoint has a bad minute. Components, and the three metrics that page you?
The endpoint does exactly three things inline: verify the HMAC signature, write the raw event to a durable queue, return 200 — a few milliseconds, so provider timeouts (and thus retry storms) mostly can't start. Workers consume the queue, check each event_id against a processed-IDs table (insert-if-absent, atomic with the side effect), skip duplicates, and treat events as doorbells — fetching current state from the provider's API rather than trusting payload order. Page on: queue depth (workers falling behind — the early-warning signal), endpoint non-2xx rate (you're triggering the provider's retry schedule), and duplicate rate (a spike means acks aren't landing). The classic failure to name aloud: processing inline → slow responses → provider timeouts → retries pile onto an already-slow endpoint → a self-inflicted DDoS from your own payment provider.
Launch night. Your movie-ticket side project has run quietly for months with a few hundred users. Tonight a real theatre chain switches on, traffic climbs, and your phone starts buzzing: the "my bookings" page is timing out. You SSH in, run the query by hand, and watch it hang. Nine seconds. Nine seconds to fetch one user's twelve bookings.
The bookings table has 20 million rows now, and the database is doing the only thing it can do without help: reading every single row and asking, one by one, "does this belong to user 48213?" You type one line — CREATE INDEX ON bookings (user_id); — and the same query returns in 8 milliseconds. A thousand times faster. Nothing else changed.
That line looks like magic. Interviewers know that most engineers treat it as magic — they've typed it, seen it work, and never once pictured what the database actually built. So this is where they dig. In Amazon HLD rounds and across India interview loops especially, the schema-and-index part of a design gets graded with unusual ferocity: exact tables, exact indexes, exact locking. Hand-waving here is the fastest way to lose a level. This chapter makes sure you never hand-wave here again: how requirements become tables, why we normalize and when we deliberately stop, how indexes really work under the hood, and what a transaction actually promises you.
Data modeling starts before any SQL, and the rest of this chapter hangs off it in one chain: requirements become tables, normalization edits those tables, the queries the tables must answer become indexes, and the writes that touch several at once become transactions. Read your requirements and underline two kinds of words. The nouns become entities — the things your system remembers: users, movies, shows, seats, bookings. The verbs become queries — the questions your system must answer fast: "show me tonight's shows for this movie", "is seat F7 taken?", "list my bookings, newest first."
Entities give you tables. Verbs give you indexes and shape. A schema designed from nouns alone is a filing cabinet nobody can search; a schema designed from verbs alone is a pile of query-shaped caches with no source of truth. You need both, in that order: model the entities honestly, then check every important query has a fast path through them.
Between entities sit relationships, and there are only three shapes. One-to-one (a user has one profile). One-to-many (a user has many bookings — the "many" side carries a user_id column pointing back). Many-to-many (a user attends many shows, a show has many users — you need a third table in the middle, one row per pairing; here, that middle table is bookings itself, and it's the most important table in the system).
Suppose you're lazy and store the movie title on every booking row. Twenty million bookings, each carrying "Jawan" as text. Then the studio renames the film. Now you must update millions of rows, and if your update misses even one — a crash halfway through, a code path you forgot — your database holds two different truths about the same fact. Which one is right? Nobody knows. That's called an update anomaly, and it is the disease normalization cures.
Normalization is one idea wearing academic clothing: every fact lives in exactly one place, and everything else points at it. The title lives in movies, once. Bookings store a movie_id — a pointer, not a copy. Rename the film: one row changes, and every booking is instantly, automatically right. You'll see textbooks enumerate "normal forms" (1NF, 2NF, 3NF…); in an interview and in practice, "no fact stored twice" gets you 95% of the value, and saying it that way sounds like experience rather than a course you took.
Normalization is keeping one contacts app instead of writing your friend's address on every envelope you might someday send. When she moves, you update one entry and every future letter goes to the right place. Denormalization is pre-printing a hundred address labels because you post something to her every day and can't be bothered opening the app each time — genuinely faster, but the day she moves, you'd better remember to throw away the old labels. The analogy is honest about the cost: pre-printed labels don't update themselves, and neither do denormalized columns.
Normalized data is cheap to write and correct by construction — but some reads get expensive. "Show the like count on every post in the feed" done honestly is a COUNT(*) over a likes table, per post, per feed load. At Instagram scale that's counting millions of rows millions of times a day to display a number that changes rarely. So you break your own rule on purpose: add a like_count column on posts, increment it when a like arrives, and read it for free.
You've now stored a fact in two places (the likes rows and the counter), which means they will drift — a crash between the insert and the increment, a bug, a manual fix. Denormalizing like an adult means accepting that and planning for it. The senior formula has three parts: denormalize only when a hot read measurably can't be served the normal way; name the writer that keeps the copy fresh; and name the reconciliation story — a periodic job that recomputes the truth and repairs drift. Denormalization without a reconciliation story isn't an optimization, it's a slow-motion data-corruption project.
Every table needs a primary key — a value that uniquely names each row forever. You have two options. A natural key is a real-world attribute that happens to be unique: an email, a passport number, the pair (show_id, seat_id). A surrogate key is a meaningless identifier the database invents: an auto-incrementing bigint, a UUID.
Commit to surrogate keys as your default, and be ready to say why: natural keys look immutable until they aren't. People change emails. Countries reissue ID numbers. The moment a "permanent" natural key changes, every table that referenced it needs a coordinated multi-table update — the exact update anomaly you normalized to avoid. A surrogate key never changes because it never meant anything. Then — and this is the part that reads senior — still enforce the natural uniqueness with a unique constraint. id is the row's name; UNIQUE (email) is a business rule the database enforces at full speed with zero trust in your application code. In the bookings walkthrough below, one unique constraint turns out to be the correctness backbone of the entire system.
One practical note when someone says "just use UUIDs": random UUIDs (v4) insert into random positions in the index, wrecking the tidy append-to-the-right pattern that sequential bigints give you, which costs cache locality and write speed on big tables. If you want globally unique IDs generated across many machines, use time-ordered schemes — which is precisely why Instagram built theirs the way they did (story below).
A foreign key constraint tells the database: bookings.user_id must always point at a real row in users. The database checks it on every insert and blocks deleting a user who still has bookings. It's a free integrity guard, and for most systems, most of their lives, you should keep it — dropping foreign keys "because scale" on a system doing 200 writes/second is cargo-culting.
But large shops do drop enforcement, for three concrete reasons worth knowing. First, every FK check is extra work on every write — a lookup into the parent table, plus locking on the parent row that can create surprising contention. Second, the big online schema-migration tools that let you alter huge tables without downtime don't play well with them — GitHub's gh-ost explicitly does not support foreign keys at all. Third, foreign keys can't cross shards: once users and bookings live on different machines, there is no database left that can see both sides of the constraint. At that point the rule moves into application code plus periodic integrity-check jobs that hunt for orphaned rows. That's the honest trade: you give up an automatic guarantee and buy it back with vigilance.
GitHub has run on MySQL since the beginning — Rails app, one primary accepting all writes, replicas fanning out reads. Their schema philosophy is deliberately conservative, and their tooling shows what "careful schema at scale" costs. Altering a table with hundreds of millions of rows in MySQL was so painful — the classic approaches use triggers that add load to an already-hot primary — that in 2016 GitHub built and open-sourced gh-ost: a tool that builds a ghost copy of the table, replays changes by tailing MySQL's replication log instead of using triggers, then swaps the tables atomically. A schema change on a giant table becomes a slow, throttleable, pausable background operation instead of a scary lock. Notably, gh-ost's own documentation lists foreign keys as unsupported — one reason GitHub's schemas avoid FK constraints and enforce integrity in the application. And the "single primary" part of the story has teeth too: in October 2018, just 43 seconds of lost connectivity between their East Coast data center and the rest of the network caused automated tooling to fail writes over to the West Coast — leaving seconds of writes stranded on the old primary and putting GitHub into degraded mode for roughly 24 hours while they reconciled the two histories. A single-writer design is beautifully simple right up until the moment the "single" part is disputed.
Take the "my bookings" query against 20 million rows. Without an index, the database's only move is a full table scan — read every row, keep the matches. The fix is the same one humans invented for phone books: keep a sorted copy, and search by jumping, not reading.
A B-tree index is a phone book. Nobody finds "Sharma" by reading from page one — you open near the back, see "Patel", jump forward, see "Singh", step back, done: five glances instead of a hundred thousand entries, because sorted data lets you halve the search space with every look. And a composite index is exactly a phone book's sort order: last name, then first name. Finding "Sharma, Priya" is instant. Finding everyone named Priya regardless of surname is hopeless in the same book — Priyas are scattered across every surname. The book's order only helps if your question starts with the first column of the sort.
Concretely, a B-tree is a tree of fixed-size pages (blocks of ~8KB). The leaf pages at the bottom hold every indexed value in sorted order, each with a pointer to its actual row. The pages above are a thinning pyramid of signposts: "keys up to M are down-left, after M down-right." Because each page holds hundreds of keys, the tree is astonishingly shallow — three to four levels cover hundreds of millions of rows. A lookup walks root → branch → leaf → row: about four page reads instead of twenty million row reads — seconds of scanning become milliseconds — and the top levels are so hot they permanently live in RAM. That's the entire trick. Not magic — logarithms plus a sorted copy you paid to maintain.
An index on (city, created_at) sorts entries by city first, and by time within each city — exactly like last-name-then-first-name. So WHERE city = 'Delhi' ORDER BY created_at is a dream: jump to the Delhi block, read it in order, stop. Already sorted — no separate sort step at all. But WHERE created_at > '2026-01-01' alone gets nothing from this index: recent rows are scattered through every city block, like Priyas across surnames.
This gives you the leftmost-prefix rule — an index on (a, b, c) serves queries filtering on a, on a, b, or on a, b, c, but not on b or c alone — and one design rule worth memorizing: columns you match with equality go first; the column you range-scan or sort by goes last. Equality columns narrow you to one contiguous block; the trailing column is sorted within it, ready for your range or your ORDER BY.
Normally a lookup is two steps: find the entry in the index, then follow the pointer to the table to fetch the other columns. But if the index itself contains every column the query needs, the second step vanishes — the database answers entirely from the index. This is a covering index (Postgres calls the result an index-only scan, and lets you append extra non-searchable columns with INCLUDE). For a red-hot query like "seat map for this show," an index on (show_id, seat_id, status) means the availability check never reads the table at all. It's the closest thing SQL has to a free lunch — except the lunch is paid for on the write side, which brings us to the tax.
An index is a sorted copy that the database swears to keep synchronized forever. So every INSERT into a table with four indexes performs roughly five writes: the row itself, plus an ordered insertion into four separate B-trees — finding the right page in each, splitting pages when they fill. Updates that touch indexed columns pay twice (delete the old entry, insert the new one). Indexes are not free speed; they are a purchase of read speed using write speed as currency, plus disk, plus cache space.
The operational habit that follows: audit your indexes like subscriptions. Every database tracks per-index usage stats (Postgres: pg_stat_user_indexes). An index nobody has read in a month is a pure tax — drop it. Saying this in an interview, unprompted, is a small sentence that signals years of operating real tables.
You built the index; the database refuses to use it. Before you rage, understand that the query planner — the component that picks how to execute each query, using statistics about your data — is usually right. The common cases:
WHERE status = 'completed' matching 90% of rows: using the index means bouncing between index and table millions of times in random order. Reading the whole table sequentially is genuinely faster, and the planner knows it. Indexes shine when they eliminate most of the table, not a tenth of it.ANALYZE) — and it will make confident wrong choices from old numbers.WHERE lower(email) = '…' can't use a plain index on email — the index stores emails, not lowercased emails. You need an index on the expression itself. Same trap with type mismatches and leading-wildcard LIKE '%gmail.com' (the sorted order can't help when the prefix is unknown — Priyas again).The habit: never guess. EXPLAIN ANALYZE shows the actual plan and actual timings. Senior engineers state index designs with confidence and verify them with EXPLAIN; only one of those habits is optional.
Here's a slow page where every query is fast. Your feed code loads 50 posts (one query), then loops: fetch each post's author (50 queries), fetch each post's like count (50 more). 101 round trips. At a modest 1ms each, that's 100ms of pure back-and-forth — before any real work. This is the N+1 query problem, and ORMs (libraries that auto-generate SQL from object code) produce it silently, because a harmless-looking post.author.name inside a loop is a query.
The fix is always some form of batching: a JOIN that brings authors along with posts, or two queries total — posts, then WHERE id IN (…all 50 author ids…). ORMs call this eager loading. The senior habit is watching queries per request as a first-class metric: a request that's slow while every individual query is fast is N+1 until proven otherwise.
A transaction groups several statements into one all-or-nothing unit. The guarantees travel under the acronym ACID, and at working depth they are: Atomicity — all of it happens or none of it does; a crash mid-transfer can't debit one account without crediting the other. Consistency — your declared rules (constraints, uniqueness) hold before and after. Durability — once the database says "committed," the data is on disk and survives a power cut; the promise is made before the acknowledgment, not after. And Isolation — concurrent transactions don't trample each other. Isolation is where the interview questions live, because it isn't one guarantee — it's a dial.
Here's the reality most engineers learn the hard way: the default isolation level in Postgres (and most databases you'll touch — MySQL's InnoDB defaults slightly stronger) is read committed, and it promises only that you never see another transaction's uncommitted data. It does not stop two transactions from reading the same state and acting on it simultaneously. Watch it fail:
-- Two users, same seat, same instant. Both transactions run:
SELECT status FROM bookings WHERE show_id = 99 AND seat_id = 7;
-- both see: no booking exists. Both conclude: seat is free!
INSERT INTO bookings (user_id, show_id, seat_id) VALUES (…);
-- both inserts succeed. Seat F7 is sold twice.
Nothing "went wrong" by read committed's rules — neither transaction saw uncommitted data. The check-then-act race is simply not its job to prevent. You have two working-strength fixes. One: lock the moment of decision — SELECT … FOR UPDATE takes a row lock, so the second buyer's check waits until the first buyer's transaction commits, then sees the seat taken. Two, and better: make the database the referee — a unique constraint on (show_id, seat_id) means the second insert fails no matter what the application believed. The constraint is a seatbelt that works even when the code forgets to check; the strongest systems wear both, and treat the constraint as the backbone.
Could you instead crank the dial to serializable — the level where the database guarantees the outcome equals some one-at-a-time ordering? Yes, and it genuinely closes these races — by detecting conflicts and aborting transactions, which your code must catch and retry. Committing to "serializable everywhere" in an interview sounds rigorous but reads as inexperience: you've volunteered retry loops and throughput costs across your whole system to fix a race that lives in one query. The senior answer runs read committed globally and applies targeted force — constraints and row locks — exactly where the money is.
Somewhere in most loops, someone asks: "SQL or NoSQL here?" There's a trap in the question. The answer interviewers dread is the blog-post recitation — "SQL gives ACID and joins, NoSQL scales horizontally…" — because it's what people say when they've read comparisons instead of operating databases. The advice from the Hello Interview guides is blunt and worth internalizing: pick one database and know it deeply — broad SQL-vs-NoSQL comparisons read as inexperience. The dichotomy is mostly dead anyway: Postgres handles JSON documents, full-text search, and replicates nicely; DynamoDB has transactions. What separates candidates is depth in something: "Postgres. Here's the schema, here are the three indexes and why, here's the unique constraint doing my concurrency work, here's what I'd watch in production." Specific beats broad, every single round.
When Facebook bought Instagram in April 2012 for a billion dollars, Instagram had about 30 million users and 13 employees — and the backend was deliberately boring: Django, PostgreSQL, Memcached, Redis. Their engineering motto was literally "do the simple thing first." When one Postgres box could no longer hold the data, they didn't switch databases — they went deeper into the one they knew: several thousand logical shards (implemented as Postgres schemas) mapped onto a handful of physical machines, so future growth meant moving shards, not rewriting code. Needing unique IDs generated across all those shards without any central bottleneck, they built a 64-bit ID inside Postgres itself — a few lines of PL/pgSQL packing 41 bits of timestamp, 13 bits of shard ID, and a 10-bit per-shard sequence — after deciding that running extra ID-generation infrastructure wasn't worth the operational load for a tiny team. Years of hypergrowth, handled by one deeply-understood database and a simple schema. That's the strongest available counterexample to "you need NoSQL to scale."
Time to assemble everything into the schema you'd actually draw at the whiteboard. First the verbs — the queries that must be fast: book a seat (safely!), show my bookings newest-first, show the seat map for a show, list showtimes for a movie, log in by email. Now the schema that serves them:
CREATE TABLE users (
id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
email text NOT NULL UNIQUE, -- natural uniqueness, enforced
name text NOT NULL
);
CREATE TABLE shows (
id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
movie_id bigint NOT NULL REFERENCES movies(id),
theatre_id bigint NOT NULL REFERENCES theatres(id),
starts_at timestamptz NOT NULL
);
CREATE TABLE bookings (
id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
user_id bigint NOT NULL REFERENCES users(id),
show_id bigint NOT NULL REFERENCES shows(id),
seat_id bigint NOT NULL REFERENCES seats(id),
status text NOT NULL DEFAULT 'confirmed',
created_at timestamptz NOT NULL DEFAULT now()
);
-- THE correctness backbone: one active booking per seat per show.
CREATE UNIQUE INDEX one_active_booking
ON bookings (show_id, seat_id)
WHERE status <> 'cancelled';
-- Read paths:
CREATE INDEX bookings_by_user ON bookings (user_id, created_at DESC);
CREATE INDEX shows_by_movie ON shows (movie_id, starts_at);
Every index has a name and a job — that's the standard to hold yourself to at the whiteboard:
| Query | Served by | Why it's fast |
|---|---|---|
| My bookings, newest first | (user_id, created_at DESC) | Equality on the lead column finds one block, already sorted by time — no sort step |
| Book seat F7 for show 99 | unique (show_id, seat_id) | The insert either wins or fails — concurrency handled by the constraint, not by hope |
| Seat map for a show | same unique index | One contiguous block of entries per show; add status via INCLUDE to make it covering |
| Showtimes for a movie this week | (movie_id, starts_at) | Equality first, range second — the composite-order rule applied |
| Login by email | unique (email) | A business rule and a fast path in one object |
Notice the shape of the whole chapter compressed into that table: entities from the nouns, indexes from the verbs, composite order chosen by the equality-then-range rule, uniqueness doing the concurrency work, and not one index without a paying customer. There's even a partial index up there — the WHERE status <> 'cancelled' clause means cancelled bookings don't occupy the seat, and the index stays smaller too. That's the schema conversation Amazon rounds want: every object justified, every hot query traced to its fast path.
Schema questions are where interviewers separate "has read about databases" from "has operated one" — and Amazon HLD rounds and India loops in particular will drill until they find the bottom of your knowledge. A passing senior answer names tables, keys, the two or three indexes with reasons, and the concurrency mechanism — unprompted. Expect these probes:
Explain to a junior engineer, in five or six sentences, what a database index is, why it makes reads fast, and why you shouldn't add ten of them.
An index is a sorted copy of one or a few columns that the database maintains next to your table, like a phone book for your data. Because it's sorted, the database can find any value by jumping — a few page reads instead of scanning millions of rows — the same way you find a surname without reading the whole book. The catch is that the database must update every index on every insert and update, forever, so each index makes reads faster by making writes a bit slower and eating disk and cache. Ten indexes means every write updates eleven structures — and any index no query uses is pure tax with no refund. So add an index when a real query needs it, check with EXPLAIN that it's actually used, and drop the ones that sit idle.
Design the tables and indexes for a doctor-appointment app. Requirements: patients book slots with doctors; a slot can't be double-booked; patients see their upcoming appointments; doctors see their day's schedule. Write the queries first, then the schema. Ten minutes.
Queries: book a slot (safely), my upcoming appointments, doctor's schedule for a day. Tables: patients(id, email UNIQUE, name), doctors(id, name, specialty), appointments(id, patient_id, doctor_id, starts_at, status). The backbone: unique index on (doctor_id, starts_at) where status is active — the database referees double-booking, whatever the app code does. Read paths: (patient_id, starts_at) for "my upcoming" (equality then range — the composite rule), and the unique index already serves the doctor's-day query since equality on doctor_id plus a range on starts_at is its exact leftmost-prefix shape. Three indexes total, each with a paying query, and concurrency handled by a constraint. If you wrote more indexes than queries, revisit the write tax.
You have an index on (customer_id, order_date). For each query, will the index help? (a) WHERE customer_id = 7 AND order_date > '2026-01-01'; (b) WHERE order_date > '2026-01-01'; (c) WHERE customer_id = 7 ORDER BY order_date DESC LIMIT 10; (d) WHERE upper(customer_id::text) = '7'.
(a) Yes — the ideal shape: equality on the lead column jumps to one block, then a range over the sorted trailing column. (b) No — without the leading column, matching dates are scattered through every customer's block (the Priya problem); the planner will scan. (c) Yes, beautifully — equality plus the trailing column's sort order means the database reads 10 entries from the block's end and stops; no sort, almost no work. This is how you paginate cheaply. (d) No — you wrapped the column in a function, so the stored sorted values no longer match what you're comparing; the index can't be used unless you index that expression itself.
Your bookings system checks seat availability with a SELECT, then inserts if free. It runs at read committed. A load test with 200 concurrent buyers on one hot show sells several seats twice. Explain the exact interleaving that causes this, and give two fixes with their trade-offs.
Interleaving: buyer A's SELECT sees no booking for seat F7; before A's INSERT commits, buyer B's SELECT runs and also sees no booking (A's insert is uncommitted, and read committed hides uncommitted data — which is exactly its job); both INSERTs then succeed. No isolation rule was violated — check-then-act is simply outside read committed's guarantees. Fix 1: unique constraint on (show_id, seat_id) for active bookings — B's insert now fails with a constraint error your code turns into "seat taken." Cheap, bulletproof, works even when future code forgets to check; the cost is handling the error path. Fix 2: SELECT … FOR UPDATE on the seat row — serializes buyers at the decision point, giving a cleaner "check saw the truth" flow; the cost is lock waits piling up on hot shows. Best answer: both — the lock for flow, the constraint as the seatbelt.
A teammate proposes adding seats_left to the shows table so the listings page stops running COUNT(*) over bookings. Argue for it like a senior: when is it justified, and what three things must the design include?
Justified when the read is measurably hot (listings render on every visit; the COUNT scans a hot table thousands of times a minute) and the normalized path can't be made cheap enough with an index alone — check that first, since a covering index on (show_id, status) might make COUNT nearly free. If you proceed: (1) a named writer — decrement in the same transaction as the booking insert, so atomicity keeps counter and truth aligned on the happy path; (2) a reconciliation job — periodically recompute from bookings and repair drift, because drift is a when, not an if; (3) an honest tolerance — the listings count may be seconds stale, and that's fine, but the actual booking decision still goes through the unique constraint, never through the counter. Denormalized values are for display; constraints are for truth.
Twenty minutes into a senior loop, a candidate says the magic words: "I'll shard the database." The interviewer, who has been nodding pleasantly, asks one quiet question: "How much data is there?" Silence. The candidate has 50 million users in the requirements and no idea what that means in gigabytes. They do the math live, badly, and it comes out to about 40 GB. Forty gigabytes. That fits in the RAM of one mid-sized server. They had just proposed slicing a suitcase into six pieces and shipping each piece on its own truck.
Nothing about that candidate's knowledge was wrong. What failed was the ritual that should run before any architecture word leaves your mouth: a sixty-second envelope calculation that tells you what size of problem you're actually holding. Skip it, and every decision after is a guess wearing a confident voice.
The ritual is small: a handful of memorized numbers, a few conversion tricks, and one iron rule. Learn it and you can take "10 million daily users" and, in a minute of talking out loud, know whether you need one server or one thousand — and, more importantly, what that answer changes about your design.
Here is the rule, up front, because everything else hangs on it: an estimate that doesn't change your design is filler. Interviewers watch candidates compute "36 terabytes per year!" with visible pride, write it in a corner of the whiteboard, and then design exactly the system they were always going to design. That's not estimation. That's a magic trick with no reveal.
Every calculation must end with the word "so." Reads outnumber writes 200 to 1 — so I design the read path first and lean on caching. The dataset is 40 GB — so it fits in memory and sharding is off the table. Media is 95% of the bytes — so it goes to object storage and a CDN, and the database only holds metadata. The number is the setup; the "so" is the deliverable. If you can't finish the sentence, don't do the math.
Google's SRE (Site Reliability Engineering) organization has a name for this discipline: Non-Abstract Large System Design, or NALSD — and it built an interview format around it, documented in The Site Reliability Workbook (2018). The format is an iterative squeeze: the candidate sketches a basic design; the interviewer asks Is it possible? Can we do better? Then the scale goes up and the questions sharpen: Is it feasible with real hardware and a real budget? Is it resilient when machines die? The workbook walks a full example — joining two streams of ad logs — and judges every iteration purely on whether the numbers close: disk seeks per second per machine, machines implied, whether the answer costs absurd money. The reason for "non-abstract" is blunt: a design that exists only as boxes and arrows cannot be evaluated at all. Google made "convert the boxes into machines, disks, and dollars" a formal hiring bar because engineers who can't do it ship designs that physically cannot work — and nobody finds out until the hardware order arrives.
Envelope math is how you shop for a wedding venue. You don't count guests to the exact person — you say "roughly 300 people, so the 80-seat hall is out, the stadium is absurd, we want a mid-sized banquet hall." The estimate is crude, and it doesn't matter: its only job was to eliminate categories. That's the goal of interview estimation too — not the right number, but the right category of solution. The analogy breaks in one place: venues forgive a 10% overflow; a database at 110% of capacity does not.
You can't estimate without anchors. The anchors are a short table of latencies — how long basic operations take — passed around the industry for decades. Memorize it like multiplication tables: not because anyone quizzes the raw values, but because every storage and caching decision you'll ever defend is secretly a ratio of two rows in it.
| Operation | Cost | Scaled: if L1 took 1 second |
|---|---|---|
| L1 cache reference | 0.5 ns | 1 second |
| Main memory (RAM) read | 100 ns | ~3 minutes |
| Send 1 KB over a 1 Gbps link | 10 µs | ~6 hours |
| SSD random read | 100 µs | ~2.5 days |
| Read 1 MB sequentially from RAM | 250 µs | ~6 days |
| Round trip inside one datacenter | 0.5 ms | ~11 days |
| Read 1 MB sequentially from SSD | 1 ms | ~23 days |
| Spinning-disk seek | 10 ms | ~8 months |
| Round trip across an ocean | 150 ms | ~9.5 years |
Read the third column slowly. If touching the CPU's own cache is one second, reading RAM is a stroll to the next room, an SSD read is a weekend trip, a disk seek is most of a year, and asking a server on another continent is a decade. These aren't differences you tune around — they're different universes, and the ratios between them are the actual content of system design:
The table has a documented lineage. Peter Norvig published an early version in his 2001 essay "Teach Yourself Programming in Ten Years," as approximate timings every programmer should internalize. Then Jeff Dean — the Google engineer behind MapReduce, Bigtable, and Spanner — put a slide titled "Numbers Everyone Should Know" into his 2009 LADIS keynote, and it went viral the way engineering slides occasionally do. Dean's point was never trivia: he and Sanjay Ghemawat were famous inside Google for refusing to build anything until the envelope math closed. In his talks, Dean estimates rendering an image-search page of 30 thumbnails — read them serially from disk at ~10 ms a seek and it's over half a second; read them in parallel and it's tens of milliseconds — the entire architecture choice falling out of arithmetic you can do in your head. A researcher named Colin Scott later built a version of the table projected across the years, revealing the deeper lesson: SSDs and networks got dramatically faster since 2009, but the cross-ocean row barely moved. Silicon improves; the speed of light doesn't.
Second tool: fluency with big numbers, so the arithmetic never steals attention from the reasoning. Three habits cover it.
Habit one: powers of two collapse into powers of ten. 2¹⁰ = 1,024 ≈ 1,000. So kilo ≈ 10³, mega ≈ 10⁶, giga ≈ 10⁹, tera ≈ 10¹², peta ≈ 10¹⁵, and you can stop caring about the 2.4% error forever. Attach a physical anchor to each: an MP3 song ~5 MB, a two-hour HD movie stream ~4 GB, a laptop SSD ~1 TB — and the full text of English Wikipedia is only ~20 GB compressed, exactly the kind of anchor that stops you from claiming a text dataset is "petabytes."
Habit two: know what things weigh.
| Thing | Size | Worth remembering because |
|---|---|---|
| ASCII character | 1 byte | English text is cheap |
| UTF-8 character | 1–4 bytes | Emoji and most non-Latin scripts cost more; budget 2× for global text |
| int / long / timestamp | 4 / 8 / 8 bytes | IDs and times are nearly free |
| UUID | 16 bytes (36 as text) | Store it as bytes, not a string |
| A tweet-sized text row | ~300 bytes raw → budget 1 KB | Indexes, row overhead, and metadata multiply raw data 2–3× |
| Thumbnail image | ~10 KB | Small enough to consider caching aggressively |
| Compressed photo | ~200 KB | What you actually serve |
| Phone-camera original | ~3 MB | What users actually upload — 15× bigger than what you serve |
| One minute of HD video, streamed | ~40 MB | Why video companies talk about petabytes and everyone else doesn't |
Notice the built-in multipliers: real storage is raw bytes × 2–3 for indexes and overhead, then × 3 for replication (three copies, so a dead disk loses nothing). A "10 GB" dataset occupies more like 60–90 GB of actual disk. Say the multipliers out loud; forgetting replication is the classic silent 3× error.
Habit three: round ruthlessly, and announce it. One significant figure everywhere. A day has 86,400 seconds — call it 10⁵ and accept the 14% error with a smile. This buys the most useful conversion in the discipline: 1 million events per day ≈ 10 per second; 1 billion per day ≈ 10,000 per second. Chant it. A candidate computing "23.148 requests per second" to three decimals is broadcasting that they think precision is the point. It isn't — being reliably within 3× of reality, fast, is the point, and saying "I'm rounding a day to 100K seconds" shows your roundings are choices, not accidents.
Now the workhorse calculation, the one you'll run in every interview: converting a population of humans into a rate of requests — QPS, queries per second — in four hops.
Three details make this senior-grade instead of mechanical:
Use DAU — daily active users — not registered users. "500 million users" in a prompt usually means accounts; maybe 20% show up on a given day. Asking "registered or daily active?" is a free display of experience — the ratio can be 5–10×, the difference between a rack and a fleet.
The peak factor is a claim about human behavior, so justify it. Consumer traffic is a tide: near-zero at 4 a.m., cresting in the evening. For a typical diurnal (daily-cycle) pattern, peak hour runs 2–5× the daily average — pick a number and say why: "consumer app, single dominant country, evening peak — I'll design for 3×." A global user base flattens toward 2×; anything event-driven (sports, TV, flash sales) can spike far past 5×, which we'll get to shortly.
Extract the read/write ratio — it names your architecture. Run the pipeline twice, reads and writes, and divide. A social feed comes out 100:1 or steeper: design the read path first, cache aggressively, and let writes do extra work (like precomputing feeds) to keep reads cheap. A logging pipeline is the mirror image, 1000:1 write-heavy: now batching, append-only storage, and sequential throughput matter, and read latency is nearly irrelevant. One division tells you which of two very different books you're designing from.
QPS means nothing until divided by what one machine can do. Carry these anchors:
Then the formula: servers = peak QPS ÷ per-box capacity, × 2–3 for headroom. The headroom isn't cowardice — it covers deploys (some boxes are always out of rotation), failures (design for N−1: the fleet must carry peak with one box dead), and the gap between benchmark and reality. And when the formula outputs "2.4 servers," believe it: the correct reading is scale is not this system's problem, and the minutes you were about to spend on sharding belong somewhere else.
For years Stack Overflow's engineers — most publicly Nick Craver, in his architecture write-ups around 2016 — documented a fact that quietly humiliates over-engineered designs everywhere: one of the largest programming sites on the internet, serving on the order of 200 million HTTP requests a day, ran its web tier on about nine servers idling at roughly 5–10% CPU. Craver was explicit that the site could run on far fewer; nine existed for redundancy and rolling deploys, not capacity. The SQL Servers behind it held hundreds of gigabytes of RAM so essentially all hot data was served from memory — the latency table in action. Run the envelope math yourself: 200M ÷ 100K seconds = ~2,000/second average; spread over nine boxes, a couple hundred per box, loafing. The lesson isn't "you only need nine servers" — their workload is unusually read-heavy and cache-friendly. The lesson is that they knew their numbers cold, sized the system to reality instead of fashion, and spent the complexity budget where it paid: keeping the hot set in RAM.
Two conversions cover almost every bandwidth question. First, network links are quoted in bits, storage in bytes — divide by 8. A 1 Gbps link moves 125 MB/s; a "10 Gbps NIC" (the server's network card) means ~1.25 GB/s; mixing these up is a silent 8× error. Second, ingress and egress are wildly asymmetric for anything media-shaped: users upload a photo once and the world views it a thousand times, so egress (data out) routinely dwarfs ingress 10–100×. Egress is also what clouds charge for — which is why every media-heavy design routes views through a CDN (a network of edge servers caching content near users) and reserves origin bandwidth for cache fills.
Real access patterns are savagely skewed: a small fraction of items gets almost all the traffic. The shorthand is the 80/20 rule — 20% of items receive 80% of requests — and real popularity distributions (Zipf curves, where the k-th most popular item gets roughly 1/k of the top item's traffic) are usually steeper. The skew is a gift: you don't need everything in fast storage, only the hot set.
Sizing is one multiplication: hot fraction × item count × item size. And do the hit-rate arithmetic correctly: going from 90% to 99% hit rate is not "9% better" — misses drop from 10% to 1%, meaning 10× less load on the database. Miss rate, not hit rate, is what your database experiences.
A bookshop doesn't keep every book ever printed on the shop floor — it keeps this month's bestsellers on the front table and orders the rest from the warehouse on demand. The front table is a rounding error of the catalog and serves most customers instantly. Cache sizing is deciding how big the front table must be, and the 80/20 rule says: shockingly small. Where the analogy breaks: when a book goes viral, the shop just sells out — your cache instead forwards the stampede to the warehouse, which is how databases die. Hold that thought for the caching chapter.
The ritual end-to-end, three times. Each finishes with the only question that matters: so the design changes how?
Say 200M DAU. Posting is rare: one post per two users per day → 100M posts/day → 1K writes/second average, 3–5K peak. Reading is constant: ~100 posts scrolled per user per day → 20B reads/day → 200K/second average, 600K+ peak. Ratio: 200:1 read-heavy. Storage: ~300 bytes of text per post, budget 1 KB with indexes and metadata → 100 GB/day → ~36 TB/year before replication. Media: if 10% of posts carry a 200 KB image, that's 2 TB/day — media outweighs five years of text in under three months.
So the design changes how? Three ways. The 200:1 ratio means the read path is the product: precompute feeds at write time (spend the cheap resource, writes, to subsidize the expensive one, reads) and cache everything. The write rate — a few thousand per second — is within reach of a relational database, so sharding is motivated by the 36 TB/year of data volume, not write QPS; say that distinction out loud, it lands. And media never touches the database: object storage plus CDN, with the DB holding only URLs.
A photo app, 50M DAU, heavy uploaders: 20M photos/day at ~3 MB per phone original. Ingress: 60 TB/day ÷ 100K sec ≈ 600 MB/s ≈ 5 Gbps average, ~15 Gbps peak. The other direction: each DAU views 50 photos/day served compressed at 200 KB → 500 TB/day ≈ 46 Gbps average egress, well over 100 Gbps at peak. Egress is ~10× ingress — after the 15× compression from originals to served copies.
So the design changes how? First: 15 Gbps of uploads must not flow through your app servers — that's saturating NICs with bytes your application logic never touches. Clients upload straight to object storage using presigned URLs (temporary permission slips the app server hands out); the app server handles only metadata. Second: a queue-driven pipeline compresses each 3 MB original asynchronously, because that 15× reduction is what makes the egress bill survivable. Third: 100+ Gbps of peak egress is a CDN's job by definition; your origin serves cache fills only. Three architecture decisions, all falling straight out of arithmetic.
Back to the Twitter-like service: can we cache our way out of 600K reads/second? What's hot is recent: ~100M posts from the last day at 1 KB is 100 GB; the 80/20 hot set is ~20 GB. Add per-user precomputed feed lists — 200 post IDs × 8 bytes ≈ 1.6 KB × 200M users ≈ 320 GB. Total ~350 GB: by capacity, two or three big-memory boxes. But check the other axis: 600K ops/second ÷ 100K per cache node = 6–8 nodes minimum for throughput, before replicas.
So the design changes how? The happy conclusion: the entire hot working set fits in RAM at trivial cost, so cache-first isn't an optimization, it's the architecture — the database is demoted to misses and history. The subtle conclusion: this cluster is sized by ops per second, not gigabytes — you'd run ~8 nodes even if the data were 10× smaller. Junior candidates size caches in GB; seniors check both axes and name which one binds. And each node gets a replica, for a reason the arithmetic makes vivid: losing one node sends its share of 600K/second — over 70K reads — at a database provisioned for a whisper of that.
Everything so far divided totals by 86,400 as if traffic fell like steady rain. It doesn't — and here estimation stops being arithmetic and becomes judgment. Load clusters along three axes, and the cluster, never the average, is what breaks you.
Temporal clustering. The daily tide gives you the 2–5× factor. But scheduled events break the tide model entirely: a World Cup final, a flash sale, a New Year's midnight. WhatsApp has reported handling on the order of 75 billion messages on a single New Year's Eve — the world's midnights arriving as vertical spikes, time zone by time zone, each compressing an hour of traffic into minutes. Event-shaped usage means your peak factor isn't 3×; it can be 10× or worse, arriving in seconds — too fast for autoscaling to boot machines. You pre-provision for the calendar, or architect the spike away (queues, staggered fanout), or you make the news.
Geographic clustering. A "global" service with 70% of its users in one country doesn't experience its global average — it experiences that country's evening, concentrated in one region, while servers elsewhere doze. Divide by region before dividing by seconds. And the latency table looms: those users sit 150 ms from a wrong-continent deployment, and no code optimization gets that back.
Entity clustering. Per-user averages hide the celebrity with 100M followers and the one auction item everyone wants. Per-entity skew creates hot keys and hot partitions — the crux of several later chapters. For now, the habit: after computing any average, ask "and what does the worst single user/key/row look like?" The answer is routinely 10,000× the mean.
This is the line between junior and senior estimation. Junior estimation is uniform: totals divided by 86,400, users interchangeable, the world a smooth paste. Senior estimation is lumpy: it asks where traffic lands, when, and on which key — because capacity planning against the average is planning to fail at the exact moment everyone is watching.
The interviewer doesn't grade your arithmetic — they grade the connection between your numbers and your boxes. A passing answer computes fast, rounds shamelessly, states assumptions, and lands every number in a design decision. Probes to expect:
A junior teammate asks: "Why do you always scribble numbers before drawing boxes? The numbers are all made up anyway." Answer them in five or six sentences, including how you'd turn "20M daily users, 10 app opens each" into a server count.
The numbers are made up, but only to within 3× — and that's enough to pick between designs that differ by 1,000×. Watch: 20M users × 10 opens = 200M requests/day; a day is ~100K seconds, so 2K/second average; evening peak maybe 3×, so 6K/second. One app server handles 10K+ simple requests/second, so with redundancy and headroom this is three or four servers — a shelf, not a datacenter. That minute of scribbling just told us scale isn't the hard problem, so we won't waste a week designing sharding we'll never need — we'll go find the actual hard problem. That's the whole trick: the estimate isn't for knowing the number, it's for choosing what to build. An estimate that doesn't change a decision wasn't worth doing.
URL shortener: 100M new short links/month, each read 10×/month on average, stored forever. Estimate write QPS, read QPS, and five-year storage (a stored link ≈ 500 bytes with indexes). Then the real question: what did the numbers just tell you about where the hard part is?
Writes: 100M/month ≈ 3.3M/day ≈ 33/second — round to 40, peak maybe 150. Reads: 10× that — 400/second average, ~1.5K peak. Storage: 100M × 12 × 5 = 6B links × 500 B = 3 TB. Every number is tiny: one database with a replica and a small cache yawns at this, and 3 TB fits on one disk. So the crux was never scale — it's generating billions of unique short IDs without collisions across concurrent servers, and the estimate's job was to redirect your 45 minutes toward that. Candidates who skip the math spend twenty minutes sharding a 3 TB dataset; the numbers would have told them not to.
Your session store holds one entry per logged-in user: 30M concurrent sessions at peak, ~2 KB each. The ops team proposes a 16-node sharded cluster. Run the numbers and give a verdict.
30M × 2 KB = 60 GB — fits comfortably in one modern cache node's RAM (128–256 GB boxes are routine). Check the other axis: the store is hit once per user request, and 30M concurrent sessions is not 30M people actively clicking — most logged-in users are idle at any instant. If the app peaks at, say, 60K requests/second (each session averaging a request every eight minutes), that's 60K ops/second — inside one Redis node's ~100K, snug enough that you'd state the assumption out loud. So 16 shards is over-engineering by an order of magnitude; the right shape is one primary plus one or two replicas, where replicas exist for availability (sessions vanishing = everyone logged out = incident), not capacity. Senior verdict in one sentence: "capacity says one box, throughput says one box, so the only reason for more boxes is failover — two or three, not sixteen."
A cricket app sends a push notification — "Match starting!" — to 50M users, and history shows 20% open the app within one minute. Your normal peak is 8K QPS, provisioned with 3× headroom. What happens, and what do you change?
10M app opens in 60 seconds ≈ 170K QPS — 20× your normal peak, 7× your provisioned ceiling. The frontend collapses, retries multiply the load, and the outage outlasts the spike. Autoscaling won't save you: machines take minutes to boot; this spike arrives in seconds. Fixes, in order of leverage: you control the spike because you sent the notification — stagger the send over 5–10 minutes (170K becomes ~20K); pre-scale before scheduled matches, since the calendar is known; make the app-open path cheap (match screen from CDN/cache, defer personalization); rate-limit at the edge so overload degrades gracefully. The lesson: a self-inflicted synchronized event makes "peak = 2–5× average" catastrophically wrong — the distribution, not the average, is the design input.
Your service runs only in Virginia, USA. A user in Mumbai loads a page that makes 4 sequential API calls (each depends on the previous one's result). Using only the latency table, estimate their wait — and say which fixes work and which don't.
Cross-continent round trip ≈ 150 ms; India to US East is at least that — call it 200 ms real-world. Four sequential calls = 4 round trips = ~800 ms of pure travel time, before the first-connection tax (TCP + TLS handshakes cost 2–3 extra round trips) and before your servers do any work. Your code could run in zero time and the user still waits 800 ms, because the delay is the speed of light in fiber. Fixes that work: collapse the 4 calls into 1 batched call (800 → 200 ms — the cheapest fix by far); serve cacheable pieces from a CDN edge near Mumbai; eventually, deploy a region in India. Fixes that don't: faster servers, more servers, code optimization. The senior habit: count round trips first — at intercontinental distance the network isn't part of the latency budget, it is the latency budget.
Your side project got picked up by a newsletter. Yesterday you had 200 users; today the graph looks like a wall. Your one server — a modest VM running the app and the database — is at 96% CPU. Requests that took 80 milliseconds now take 4 seconds. Some just time out.
Every scaling story on earth starts at this moment, and there are only two doors out of the room. Door one: get a bigger server. Door two: get more servers. Every architecture you'll ever draw is some combination of those two moves.
Most engineers have been trained to sneer at door one. That training is wrong, and interviewers know it. So before we build the fleet, the big single box gets the honest defense it deserves.
Vertical scaling means making one machine stronger: more CPU cores, more RAM, faster disks. Horizontal scaling means adding machines and spreading the work. The industry spent fifteen years romanticizing the second and forgetting how absurdly powerful the first has become.
Here's one machine in the mid-2020s: cloud instances rent with hundreds of CPU cores and terabytes of RAM — the biggest memory-optimized ones pass 20 TB. A single NVMe SSD does on the order of a million random reads per second, and a server holds a stack of them. Run the numbers from the estimation chapter: a service doing a few thousand requests per second with a working set of a few hundred gigabytes fits on one machine with room to spare. That's not a toy — that's a successful business.
Vertical scaling also has a property no benchmark shows: it adds zero new problems. One machine has no network partitions inside it, no distributed transactions, no "which server has the session" puzzle. Every problem in the rest of this book is one you buy the moment you go from one machine to two. A senior engineer delays that purchase as long as the numbers allow.
For most of the 2010s, Stack Overflow — about 66 million page loads a day in its published 2016 numbers — ran on-premises on hardware you could inventory on one whiteboard. The well-documented 2016 snapshot: nine IIS web servers running the whole application, idling at 5–10% CPU. The databases were four SQL Server machines stuffed with RAM — 768 GB on the biggest — so hot data lived in memory and rarely touched disk. In front: a small set of HAProxy load balancers in redundant pairs. No microservices, no NoSQL fleet. Their engineers wrote openly that the site could survive on a fraction of that hardware; the rest was redundancy and headroom. The monolith only began moving to Azure in 2023, starting with the Teams product — a business decision, not a scaling emergency. The lesson isn't "never scale out"; it's that scale-up plus discipline goes much further than conference talks admit, and saying when one big box is enough earns more trust than reflexively drawing twelve services.
So why ever leave? Two reasons, and only one is performance.
So you buy a second server. And before it serves a single request, you have two new problems: traffic has to find it, and — nastier — your app has to survive users bouncing between machines. The nasty one first, because until it's solved, the second server is useless.
Run the same app on two servers and here's day one. Priya logs in; the request lands on server A, which stores "session 4f2a = Priya, logged in" in its own memory and hands back a session cookie. Her next click lands on server B, which looks up session 4f2a in its memory, finds nothing, and shows her the login page. Every second click, she's a stranger.
The bug isn't the second server. It's that server A was stateful: it kept information between requests in its own memory. The fix is to make every app server stateless — after finishing a request, the server remembers nothing, so any server can handle any request from any user. The state doesn't disappear; it moves to one of two places:
The same rule covers everything servers quietly hoard: uploaded files go to blob storage (chapter 1.7), local caches become a shared cache tier (chapter 1.6), scheduled jobs move to a queue (chapter 1.11). Interchangeable servers are cattle, not pets: add ten, kill three, and no user can tell.
A stateful server is your neighborhood barber: he knows your usual cut — in his head. Wonderful, until he's sick or has a queue of nine, and nobody else can serve you. A stateless server is a bank teller: your account lives in the central system, so any window can serve you, the bank can open five more at noon rush, and a teller going to lunch loses nothing. The honest edge: the tellers now all depend on that central system — you haven't deleted the state, you've relocated it, and the new home must itself be made reliable.
Servers are now interchangeable. Next: someone has to stand at the door and deal the requests out.
A load balancer (LB) owns the public address of your system and forwards each incoming connection or request to one of the servers behind it. Clients only ever see the LB; the fleet behind it grows, shrinks, and churns invisibly. But "forwards traffic" hides the key split: at what layer does the balancer look? The names come from the OSI model of chapter 1.1 — layer 4 is transport (TCP/UDP), layer 7 is application (HTTP).
/api/* to one pool and /video/* to another, retry a failed request elsewhere, compress, rate-limit, set cookies. Cost: real work per request, so a lower ceiling per box. Examples: NGINX, HAProxy, Envoy, AWS Application Load Balancer.An L4 balancer is the postal sorting facility: it reads the address on the envelope and routes it — millions per hour, never opening a single letter. An L7 balancer is the office mailroom clerk: she opens each envelope, reads "this is an invoice," and walks it to accounting; complaints go to support; letters from a known pest get refused (rate limiting). She's far smarter and far slower per item — which is why the postal service doesn't open mail, and why the biggest sites put L4 in front of L7.
The interview default: say "load balancer," mean L7, and reach for L4 when the numbers demand it — millions of long-lived connections, or the front door of a CDN-scale system.
The LB has a request and eight healthy servers. Which one gets it?
One senior-level refinement: power of two choices. Tracking the exact least-loaded server across a huge fleet is expensive, and when many LBs all chase the same briefly-idle server, they dogpile it. Instead: pick two servers at random, send to the less loaded. Sounds like a party trick; provably close to optimal. Envoy and NGINX both implement it.
| Algorithm | Reach for it when | Watch out for |
|---|---|---|
| Round robin | Identical servers, short uniform requests | Slow requests pile up on unlucky servers |
| Weighted round robin | Mixed hardware; canary deploys | Weights go stale as hardware changes |
| Least connections | Request duration varies a lot | Herding when many LBs chase one "idle" server |
| IP hash | Same-client-same-server without cookies | NAT lumps thousands of users onto one server |
All four share one assumption so obvious it's easy to miss: that the servers in the rotation are alive. What happens when one isn't?
Server 5 just died. Round robin doesn't know — it cheerfully deals every eighth request to a corpse, and one in eight users gets an error until someone notices. The LB needs to know who's healthy, two complementary ways:
GET /healthz every 5 seconds, mark down after 3 straight failures, up after 2 successes. Do the arithmetic aloud: 3 × 5s means a dead server keeps taking traffic for up to ~15 seconds. Tighter intervals detect faster but tolerate less jitter — naming that knob is a senior tell.Now the trap that has caused real outages: what should the health check check? The tempting answer is everything — have /healthz ping the database, the cache, the downstream services. Walk through a 10-second database blip: every server's deep check fails simultaneously, the LB marks the entire fleet unhealthy, and users get nothing — for a blip the app might have partially survived (cached reads still worked!). You converted a degradation into a total outage.
The senior answer: keep the LB's check shallow — "is this process up and able to serve?" — and handle dependency failures in the request path with timeouts and fallbacks (chapter 2.7). Kubernetes formalizes the split as liveness ("restart me") vs readiness ("don't send traffic right now"). And use the escape hatch good LBs ship: fail-open — if more than some fraction of the fleet looks unhealthy (Envoy calls it the panic threshold), assume the check is lying and route anyway. Degraded service beats no service.
Health checks also give you civilized maintenance. Kill server 3's process to deploy, and every in-flight request on it dies mid-response. Instead, connection draining: tell the LB to stop sending it new requests, let in-flight ones finish — with a deadline, typically 30–300 seconds — then deploy, then let the health check readmit it. The standard trick: the app deliberately fails its own health check (a "drain me" switch), waits out the drain, deploys. Zero dropped requests, every deploy.
The mirror image: a freshly added server has cold caches and empty connection pools. Hand it a full traffic share in its first second and it's your slowest server for a minute — or tips over. LBs offer slow start: ramp a new server from a trickle to full share over a window. Drain on the way out, ramp on the way in — that pair is what "zero-downtime deploy" means mechanically.
Back up. When we hit the session problem, there was a third fix we skipped: leave the servers stateful and make the LB keep each user on the same one. That's a sticky session (session affinity): the L7 balancer sets a cookie recording which backend served you and routes you back there every time. No Redis, no token redesign, one config flag.
Genuinely tempting — and it silently un-solves everything this chapter solved. Load skews: heavy users stuck to server C stay stuck; no algorithm can rebalance them. Deploys turn hostile: draining C means waiting hours for sessions to end, or logging its users out. Autoscaling feeds only new users to new servers. A crash wipes every session on the box. Stickiness is statefulness wearing a load balancer costume.
When is it legitimate? Two honest cases. Long-lived connections — WebSockets are pinned to one server for the connection's life by nature (rebalancing them after a restart is a chapter 3.4 problem). And legacy apps you can't rewrite yet — carts in memory, scale needed now; stickiness is the pragmatic bridge. The senior framing: stickiness as a performance optimization (warmer local cache, degrading gracefully to a miss elsewhere) is acceptable; stickiness as a correctness requirement (wrong server = broken app) is debt with a payoff date.
Look at your diagram. You removed the single point of failure by adding five servers… then put a single load balancer in front. If it dies, five healthy servers serve nobody. You moved the SPOF, not removed it. "Your LB is now the SPOF — what do you do?" is one of the most reliable senior probes in the canon, and the real answer comes in layers.
Layer one: an active-passive pair with a floating IP. Run two identical LBs sharing one virtual IP (VIP) — the address clients connect to, owned by whichever box is active. The pair runs VRRP (Virtual Router Redundancy Protocol; keepalived is the common implementation): the active LB sends a heartbeat roughly every second. Miss about three, and the standby declares it dead and claims the VIP — broadcasting "this IP now lives at my hardware address" (a gratuitous ARP). Traffic flows to the standby within seconds. Clients see at most a blip of reset connections; no DNS change, no client cooperation.
The standby LB is a theater understudy: she watches every show from the wings (heartbeats), wearing the same costume (same VIP, same config). The night the lead collapses mid-act, she steps out and continues the scene, and the audience — who bought tickets to the character — mostly never notices. Honest edges: the scene stutters for a beat (in-flight connections on the dead LB are lost), and an understudy protects against one collapse, not the theater burning down. For that you need another theater — which is where DNS and anycast come in.
Layer two: DNS failover, for when the whole site dies. A VIP pair protects a machine, not a building — a data center power cut kills both LBs. DNS failover health-checks your sites from outside (as Route 53 does) and changes the DNS answer to point at the survivor. It works, but respect its speed limit: DNS answers are cached by resolvers and operating systems for the TTL (time-to-live) you set — commonly 30–60 seconds in failover setups — and plenty of resolvers and long-lived apps ignore your TTL and hold stale answers for minutes or hours. DNS failover is measured in minutes: disaster recovery, not high availability.
Layer three: anycast, the Cloudflare-style front door. There's a way to fail over globally without DNS: announce the same IP address from many places at once. Internet routing (BGP — the protocol networks use to tell each other which addresses they can reach) delivers each packet to the topologically nearest site announcing it. That's anycast. Cloudflare announces the same addresses from data centers in hundreds of cities: a user in Mumbai reaches the Mumbai edge, a user in Lyon reaches Paris — same IP. When a site dies, it stops announcing, and the internet reroutes its users to the next-nearest edge in seconds, with no DNS change. That's how edge networks put a load balancer "everywhere" behind one address; the full CDN story is chapter 1.7.
Notice the shape: VRRP pair inside a site (seconds, protects a box), anycast or DNS across sites (protects a location). Naming the layers and their failover time scales separates "add another LB" from a senior answer.
Everything so far — clone the stateless server, balance across the clones — is one axis on the map interviewers carry in their heads. The scale cube (from the book The Art of Scalability) says systems grow along three independent axes:
The order is a judgment call interviewers probe: X first, nearly free once you're stateless. Z when a data tier outgrows one machine. Y when the team's shape demands it — not because a conference said so. Which brings us to the company that learned this in public.
In 2011–2012 Pinterest was among the fastest-growing sites the web had seen — past 10 million users with a single-digit engineering team. Under that pressure they adopted everything that promised infinite scale, and their famous "Scaling Pinterest" retrospective shows the September 2011 stack: MySQL plus Cassandra plus MongoDB plus Membase plus Redis plus Memcached plus Elasticsearch, with multiple sharding schemes at once. Each clustered "auto-scaling" technology, pushed to its limit, failed in its own novel way — and every novel failure was an all-hands emergency with no answers on the internet. So they reversed course and simplified brutally: sharded MySQL (4,096 logical shards at launch, each object's shard baked into its 64-bit ID), Redis, Memcached, Solr — mature tools with boring, documented failure modes. Their stated lessons became canon: keep it simple; use known, proven technology; clone stateless web servers behind load balancers for traffic (the X-axis) and shard data where you must (the Z-axis). The architecture that carried Pinterest through hypergrowth was less exotic than what many ten-person startups run today.
Load balancing is rarely the headline question; it's the terrain under every question. A passing senior answer commits ("L7, least-connections because upload times vary, shallow health checks with a 15-second detection window") and raises the failure story unprompted. Probes to expect:
Explain to a junior engineer, in five or six sentences, the difference between an L4 and an L7 load balancer, and why big systems use both.
A load balancer is the front door that spreads incoming traffic across many servers, and the L4/L7 split is about how much of each message it reads. An L4 balancer only reads the envelope — IP addresses and ports — so it picks a server when the connection opens and then just forwards packets; it never knows what URL you asked for, which makes it fast enough to hold millions of connections. An L7 balancer opens the envelope: it terminates the TLS connection, reads the actual HTTP request, and can route by path or header, retry failures, and set cookies — much smarter, but it does real work per request, so each box handles less. Big systems chain them: a thin L4 layer at the front absorbs enormous connection counts, passing to L7 balancers behind it that make the intelligent routing decisions. Memory hook: L4 is the postal sorting center that never opens mail; L7 is the mailroom clerk who reads every letter and walks it to the right desk.
Your API peaks at 20,000 requests/second. Load tests show one app server handles 1,200 req/s before p99 latency degrades. How many servers behind the LB? Show the reasoning that isn't just division.
Raw division: 20,000 / 1,200 ≈ 17 servers at 100% — which you never run; a full fleet has no room for spikes, deploys, or deaths. Target ~70% peak utilization: 20,000 / (1,200 × 0.7) ≈ 24, which also survives N+2 (one server draining, one freshly dead). Commit to ~24, autoscaling between maybe 12 and 30. Senior extras: check the load test used a realistic request mix, and remember every added server multiplies queries into the database — the web tier scales while the data tier quietly becomes the next bottleneck.
What breaks: /healthz checks the app and pings the database. The database is unreachable for 20 seconds. LB config: probe every 5s, down after 2 failures, up after 3 successes. Narrate the timeline.
T+0: blip begins; every server's check starts failing. T+10s (two failed probes): the LB has marked every server down — 100% of requests now fail at the LB, including cached reads that needed no database. T+20s: database recovers, but readmission needs 3 straight successes at 5s intervals, so servers return around T+35s — a 20-second blip became a ~35-second total outage, and at readmission all the queued client retries slam the fleet at once. Fixes: shallow LB checks, dependency failures handled per-request with timeouts and fallbacks, and fail-open so that when "everyone" looks dead the LB assumes the check is lying and routes anyway.
Sketch the exact sequence for deploying new code to a 6-server fleet behind an L7 balancer with zero dropped requests, naming every mechanism from this chapter you use.
One server at a time (or two, if the capacity math allows): (1) flip server 1's drain switch so /healthz starts failing — the LB stops sending new requests (active check as removal lever); (2) wait out connection draining with a deadline, say 60s; (3) deploy; the app warms critical caches before reporting healthy; (4) after 2–3 passing probes the LB readmits it — with slow start, so it ramps instead of getting hammered cold; (5) watch error rate and p99 before touching server 2 — this is the rollback point, and weighted routing can make server 1 a 1% canary first. Repeat. Every mechanism is standard LB configuration; the skill is the sequence.
You enabled IP-hash balancing so users hit the same server (warm local caches). A week later one server sits at 90% CPU while seven idle at 15%. What happened, and what are two fixes with trade-offs?
NAT lumping: a big population — a mobile carrier's gateway, a campus — shares one public IP, so IP-hash sends the whole crowd to one server. Fix 1: cookie-based stickiness at the L7 balancer, which distinguishes users behind shared IPs; cost: sticky sessions' deploy and skew problems, though milder. Fix 2: drop affinity — least connections, with a local cache miss falling back to the shared cache tier (chapter 1.6); cost: lower hit rate, slightly higher latency, but even load and painless deploys. Say the framing aloud: affinity is fine while the "wrong" server costs only milliseconds; the moment it costs correctness or operability, externalize.
It's 7:55pm on a Tuesday, and in five minutes your e-commerce app appears in a TV ad for the first time. You've load-tested the app servers and doubled the fleet behind the load balancer. You feel ready. At 8:01 the graphs go vertical — and by 8:04 the site is timing out, because the one machine you can't clone, the database, is pinned at 100% CPU.
Afterwards you open the query log. Nothing exotic. The top query — by a mile — is the product page query. The same query, for the same two hundred trending products, thousands of times per second. Product 88871's page was assembled from scratch — four queries, some joins, a ratings average — 40,000 times in one hour. The answer changed maybe twice, when the stock count moved.
Read that again, because it's the most expensive sentence in computing: the answer rarely changed, and you computed it every single time. The database didn't fail. It was asked to do the same homework forty thousand times, and it politely did.
The fix is the highest-leverage box you will ever draw on a whiteboard: a cache. Compute the answer once, keep it somewhere fast, hand out the saved copy until it stops being true. Get this box right and most designs get easy. Get it wrong and you'll meet every failure mode in this chapter at 3am.
Start with the physics, because the numbers are the whole argument:
| Where the answer lives | Typical cost | Human scale |
|---|---|---|
| Your process's own memory | ~100 nanoseconds | Grabbing a pen off your desk |
| Redis/Memcached, same datacenter | ~0.5–1 millisecond | Walking to a colleague's desk (5,000–10,000× slower) |
| Database, well-indexed query | ~2–10 milliseconds | Going down to the archives room |
| Database, joins/scans/aggregates | 50–500+ milliseconds | Sending a clerk across town |
A cache is a bet that questions repeat. On real workloads the bet pays absurdly well, because traffic is never uniform — a small fraction of items gets most of the reads. Cache those and a little RAM absorbs most of your traffic.
Now the number that drives designs: the hit rate — the fraction of requests the cache answers without touching the database. At 10,000 requests per second (RPS) and 90% hits, the database sees 1,000 queries per second (QPS). At 99%, just 100. But flip it around, because seniors think in misses: slip from 99% to 98% — a dip that looks like rounding error — and database load doubles. The cache's job isn't "make things fast"; it's "absorb almost everything," and absorption is measured in misses. Hold that thought — it comes back with teeth at the end.
A cache is the front desk at a records office. The originals — the truth — live in the records room in the back (the database). The clerk keeps photocopies of the fifty documents people request constantly, so most visitors are served in seconds; only unusual requests send someone to the back. The catch is built in: the moment someone amends an original, the front-desk copy is quietly wrong, and the clerk has no way to know. Every hard problem in this chapter is a version of "how does the front desk find out?" — because photocopies don't update themselves.
By the time a request reaches your code, it has already passed four or five caches. Caching isn't a component you add — it's a ladder your data descends, and you choose the rungs deliberately.
Quick tour: the browser caches per user (HTTP Cache-Control). The CDN caches shared static content on servers near users (all of p1c07). A gateway can cache whole responses. An in-process "L1" map is the fastest rung but duplicated per server. This chapter's star, the distributed cache — Redis or Memcached — is a shared fleet of RAM-filled machines: computed once, a hit for every app server. Underneath, the database's buffer pool caches hot disk pages in RAM — why a "database read" is often a RAM read that still pays the parsing and network toll.
One rule organizes the ladder: the closer to the user, the faster and the staler. Each rung answers without asking the rung behind it, so each rung can be wrong for a while. Deciding how wrong for how long is the design skill.
The pattern you'll use 90% of the time is cache-aside (lazy loading): the cache is a passive bystander; your code does all the work. Every read does this dance:
value = cache.get("product:88871")
if value is None: # 1. miss
value = db.query_product(88871) # 2. compute the expensive answer
cache.set("product:88871", value, ttl=300) # 3. fill for next time
return value # hit path: one fast lookup, done
Why the default? Three reasons. It's resilient: if the cache dies, reads get slow but still work — the code falls through to the database. It's economical: only data someone asked for occupies RAM. It's universal: it works in front of anything — SQL, another service, a rendered fragment.
Its costs: the first reader after every expiry pays the full slow path, and — the dirty secret — the miss path has races. Between steps 2 and 5 the world can change: another server updates the product and deletes the key, then your step 5 fills the cache with data that was true at step 2 and is stale now — and it sits there, wrong, until the TTL saves you. Park that; it's the heart of invalidation.
Read-through is cache-aside with the dance moved inside the cache layer: on a miss the cache itself fetches from the database via a loader you configure. Same behavior, tidier code, one subtle gain — because the cache owns the miss, it can merge concurrent misses for the same key into one fetch (foreshadowing: stampedes).
Write-through changes the write path: every write goes to cache and database synchronously, so the cache is always fresh for written keys. The price: every write pays for two systems, and you cache data that may never be read. It earns its keep when read-after-write matters — a user who just edited their profile must see the edit.
Write-back (write-behind) is the fast, scary one: writes land in the cache only, flushed to the database later in batches. Writes become memory-speed, and a hundred increments to one counter flush as one row update. The terror is the gap: until the flush, the only copy lives in volatile RAM — a crash loses it. (Not exotic: your OS page cache and RAID controllers do exactly this.) The senior rule: write-back only for data you can afford to lose seconds of — view counts, likes, metrics, "last seen." Never money, never messages, never anything you promised the user you saved.
| Pattern | Who fills the cache | Write path | Main risk | Reach for it when… |
|---|---|---|---|---|
| Cache-aside | Your app, on miss | Write DB, then invalidate key | Miss races, stampedes | Default. Start here. |
| Read-through | The cache itself | Same as aside | Same, hidden in a library | You want coalescing + tidy code |
| Write-through | Every write | Cache + DB, synchronous | Slow writes, caching the unread | Read-your-own-write matters |
| Write-back | Every write | Cache now, DB later | Data loss window | Loss-tolerant, high-rate writes (counters) |
Every cached value should carry a TTL — time to live, the seconds after which the cache discards the copy and the next read recomputes it. Juniors treat TTL as a performance knob. It isn't: a TTL is a promise about maximum staleness. "Prices may be up to 5 minutes old" is a sentence a product manager can accept or veto — exactly who should weigh in. Stock tickers tolerate a second; your own profile, zero (invalidate on write instead); a "trending this week" list, an hour. Pick the TTL from the tolerance, not from vibes, and say the sentence out loud in interviews.
Then add randomness. The trap: a 9:00am deploy restarts your servers, traffic refills 50,000 keys within a minute, all with ttl=3600. At 10:00am all 50,000 expire in the same minute, and your database — idle for an hour — takes the whole site's reads at once. Synchronized filling means synchronized expiry: a perfectly scheduled, self-inflicted traffic bomb. The fix costs one line — jitter: ttl = 3600 + random(0, 360) smears the cliff across ten minutes. Add it on day one; this bug hides until the worst moment.
TTL discards what's old. Eviction discards what it must, because RAM is finite and a new value needs the space. The classic policy is LRU — Least Recently Used. Mental model: a stack of papers; every touched paper goes on top; need space, take from the bottom. It bets the recent past predicts the near future — usually true. Its blind spot: one big scan (an analytics job reading everything once) floods the top with junk and evicts your hot data. LFU — Least Frequently Used — counts touches instead, resisting scans but clinging to yesterday's viral post after tastes shift.
The honest footnote most books skip: Redis does not implement true LRU. Real LRU means bookkeeping on every read, on the hottest path. So Redis approximates: at eviction time it samples a few random keys (5 by default) and evicts the least-recently-used of the sample. Statistically close, occasionally wrong in the tail — a key you use constantly can still, rarely, be evicted. If your design assumes "this key will definitely be cached," your design is wrong; a cache entry is a hint, never a guarantee. (Redis also offers approximated LFU and random eviction.) From production: Twitter's analysis of 153 of its cache clusters (OSDI 2020) found FIFO about as good as LRU on many workloads — TTLs mattered more. Spend your whiteboard minutes accordingly.
"There are only two hard things in Computer Science: cache invalidation and naming things." — Phil Karlton
The joke lands because the problem sounds trivial — "when the data changes, update the cache" — and then you try to build it. The cache and the database are two systems with no transaction spanning them. Whatever you do, there's a window where they disagree; you only choose how wide, and at what cost. Three strategies, ascending effort:
1. TTL and accept it. Do nothing on writes; staleness is bounded by the TTL. Right more often than it feels: for a review count or trending list, 60 seconds stale costs nothing. Bounded staleness chosen on purpose isn't a bug — it's a spec. Say the bound out loud.
2. Invalidate on write. When your code writes the database, it also deletes the cache key; the next reader misses and refills fresh. Two senior details hide here. First: delete, don't set. Concurrent writers setting values can land out of order — the older value arrives last and sits there with a full TTL. Deletes can't lose that race; whoever reads next fetches current truth. Second: the cache-aside race survives — a reader that fetched just before your write can fill the cache just after your delete, resurrecting the old value. So keep a TTL backstop even when you invalidate (zombies die within minutes); fully closing the race needs the cache's help — exactly what Facebook built (below).
3. Event-driven invalidation. If many services write the same tables — or admins run SQL directly — "the app deletes on write" misses writes the app never saw. The robust version listens to the database itself: CDC, change data capture — tail the database's replication log, the stream of every committed change (tools like Debezium expose it), and turn each change into a cache delete via pub/sub. Invalidation now follows every write, whoever wrote. Cost: a pipeline to operate, and sub-second propagation delay.
And a fourth trick that sidesteps invalidation entirely: versioned keys. Put a version in the key — user:42:v7, or a deploy hash in an asset name. On change, never touch the old entry; start reading a new key, and the old one expires from disuse. No invalidation messages, immune to the races above. The standard play for config, rendered fragments, and static assets.
Facebook's worst outage in over four years — about two and a half hours down — was a cache invalidation system attacking its own database. An automated check validated a configuration value from cache; if the cached copy looked invalid, clients "fixed" it by querying the persistent store and repairing the cache. Sensible — until a change made the value in the persistent store itself invalid. Now every client, on every check, flagged the cache as bad, queried the databases, and wrote back a value that was still bad. Hundreds of thousands of repair queries per second buried the database cluster, and the loop couldn't self-heal — every fix attempt generated more queries. Engineers had to take the site offline entirely, stop all traffic to that cluster, and ease it back. The lesson: any automatic path that answers "the cache looks wrong" with "everyone go hit the source of truth" is a stampede with a mission statement. Repair paths need rate limits more than normal paths — they run exactly when the system is already in trouble.
Now the classic failure. One hot key — the homepage feed, 10,000 reads per second, TTL 60s, rebuilt by a 200ms query. The TTL expires. In the next 200 milliseconds, two thousand requests all miss and all run the expensive query simultaneously. The database, sized for a trickle of misses, gets a herd; queries slow, more pile up, and one expired key has taken down the read path. This is the cache stampede (thundering herd). Three mitigations, often combined:
1. Request coalescing — a lock on the miss. The first request to miss takes a short lock on the key (in Redis, SET lock:key NX — set if not exists) and rebuilds; the others retry the cache in a few milliseconds or — better — are served the just-expired stale value. One query instead of two thousand. This is what "leases" formalize, and what read-through libraries give you within one process.
2. Probabilistic early refresh. Never let the hot key truly expire: as the deadline nears, each reader rolls a die with rising odds; one volunteers to refresh early while everyone else keeps hitting the still-valid copy. The cliff becomes a slope. (The formal version is XFetch; the intuition is enough.)
3. Stale-while-revalidate. Make it policy: after the TTL, keep the stale value and serve it for a grace period while one background refresh runs. Users see slightly old data for seconds; the database sees one query. HTTP has this built in as the stale-while-revalidate directive — the pragmatic default wherever "5 seconds stale" is invisible, which is almost everywhere.
Facebook's 2013 paper "Scaling Memcache at Facebook" describes a memcached fleet serving billions of reads per second over trillions of items, in front of MySQL — one of the largest caching tiers ever built. At that scale, two textbook races were daily emergencies: thundering herds when a hot key was invalidated, and "stale sets" — a slow reader refilling the cache with old data just after a fresh invalidation. Their fix is the lease: on a miss, memcached hands the client a token — permission to be the one who refills the key. Other clients missing the same key get no token (one per key per 10 seconds); they briefly wait or accept slightly stale data, so the herd collapses to one database query. And a fill is only accepted if its token is still valid — any invalidation voids outstanding tokens, so a stale set is simply refused. One mechanism, both races dead. Their measurement: a workload peaking at 17,000 database queries per second without leases dropped to 1,300 with them. The lesson: at scale the miss path, not the hit path, is where caches hurt you — and the cache itself must arbitrate who may fill it.
A distributed cache spreads keys across nodes — but each individual key lives on exactly one node. Now a celebrity posts, and post:12345 draws a million reads per second. Sharding does nothing: every read goes to the one node owning the key, and no single box survives that. This is the hot key problem. Two standard treatments (p2c06 gives celebrities the full treatment): an L1 in-process cache — each app server keeps a tiny local copy of the hottest keys with a 1–2 second TTL, trading two seconds of staleness on a like count for a thousand-fold load drop. And key fan-out — write N copies (post:12345#1 … #16), each reader picks one at random; load spreads across N nodes, at the cost of N invalidations per change.
In Instagram's early years, engineers had a name for a whole class of load spike: the Justin Bieber problem. When Bieber posted, millions of likes arrived within minutes, and the naive implementation — count the like rows when someone views the post — meant re-running an expensive count over rows being hammered with concurrent writes. Co-founder Mike Krieger has said he still remembers Bieber's user ID in the database — 6860189 — because that account was so often the source of the latest fire; spikes were reliably Bieber-shaped enough that checking his account was a legitimate first diagnostic. The fix is the pattern you now know: stop recomputing — keep a denormalized, cached like count, updated incrementally, and treat ultra-popular content as its own class. Real traffic is brutally skewed; a design that handles average keys but melts on the top ten isn't done.
One miss-shaped hole remains. A request asks for deleted_user_9; the database says no such row; and since there's nothing to cache, nothing gets cached — every future request for that user hits the database, forever. A buggy client retrying a deleted resource, a crawler with an old sitemap, or an attacker probing random IDs gets a free path around your cache. The fix: negative caching — store the absence itself (user:9 → NOT_FOUND) with a short TTL, 30–60 seconds, so even the "no" is served from RAM. Short, because absence is the answer most likely to change — the user might be created a second later. DNS has cached "no such domain" since the 1990s; that's how old this trap is.
Both are RAM-backed key-value stores with sub-millisecond reads. The difference is philosophy: Memcached is a cache and refuses to be anything else; Redis is a data-structure server often used as a cache.
| Dimension | Memcached | Redis |
|---|---|---|
| Values | Opaque bytes. That's it. | Strings, hashes, lists, sets, sorted sets, streams — with atomic operations (increment, add-to-set, top-N) |
| Threading | Multithreaded — one big box uses all its cores | Single-threaded commands (I/O threads since 6.0) — scale by sharding; one slow command blocks the rest |
| Durability | None. Restart = empty. Purely a cache. | Optional snapshots (RDB) and append-only log (AOF) — can refill itself after restart |
| Replication / HA | None built in — clients shard across independent nodes | Built-in replication, Sentinel failover, Redis Cluster |
| Operational feel | Boring in the best way; little to configure or get wrong | More capable, more knobs, more ways to hurt yourself (persistence stalls, big-key blocking) |
The honest guidance: for a pure look-aside cache of flat values at enormous scale, Memcached's simplicity and multithreaded throughput are hard to beat — it's why Facebook built on it. For everyone else, Redis wins on versatility: the moment you need a leaderboard (sorted set), a rate limiter (atomic counter with expiry), a session store surviving restarts, or pub/sub for invalidation events, Redis does cache-plus-that, and one system beats two. In an interview, committing either way with the reason attached is the point; "Redis because everyone uses it" is the only losing answer.
Time to collect the debt from the hit-rate math. Your read path does 100,000 requests per second at 99% hit rate. The database sees 1,000 QPS, and that's what you've provisioned — why pay for capacity the cache makes unnecessary? Then the cache cluster goes down: a bad deploy, a partition, someone restarting it at peak "to clear a weird state." The database inherits all 100,000 QPS — a 100× overload in one second — and dies within the minute. The truth seniors design around: a cache absorbing 99% of traffic is not an optimization — it's load-bearing infrastructure, and its failure is a database failure.
So a senior design says, unprompted, what happens on that day. The playbook: degrade, shed, warm. Degrade — expensive features switch to defaults, so the core product survives on the database alone. Shed — rate-limit at the gateway so the database gets load it can survive; 30% of users served beats 100% timing out. Warm — never point full traffic at an empty cache; refill gradually, or the refill itself is a stampede. And the habits that avoid the day: cache replicas, deploy-level ceremony for restarts, and the hit rate on a dashboard with an alert — a slow decline from 99% to 96% is a silent 4× creep in database load, the outage announcing itself to anyone watching the misses.
A high-hit-rate cache is a dam upstream of a village. On a normal day the village sees only the small regulated stream the dam lets through, so everyone builds close to the banks — why not? The better the dam works, the more water it holds back, and the more the village quietly assumes it always will. The day it fails, the village doesn't get "somewhat more water" — it gets every accumulated assumption at once. Dam operators drill for failure and watch the level daily; watch your hit rate the same way. Unlike a real dam, you can rehearse the flood: test degraded mode before the failure picks the date.
Caching appears in every design round, and at the senior bar it's probed through failure, not vocabulary. A passing answer names the pattern, the TTL with its staleness contract, the stampede protection, and what happens when the cache dies — before being asked. Probes to expect:
A junior asks: "Why is cache invalidation called hard? Just update the cache when you update the database." Straighten them out in five or six sentences.
The cache and the database are two separate systems, and no transaction updates both at once — so between any two steps, another server can interleave or one system can fail. That's why we delete the cache key on write instead of setting the new value: two writers setting values can land out of order and park an old value with a fresh TTL, but a delete can't lose that race — the next reader refetches the truth. Even then, a reader that fetched old data just before your write can refill the cache just after your delete. So we keep a TTL on everything as a backstop — any zombie dies within minutes — and at Facebook scale they made the cache itself referee who may fill it (leases). The real skill isn't eliminating staleness; it's choosing and stating the bound: "at most N seconds stale, and here's the mechanism that guarantees it."
Your read path serves 20,000 requests per second through a cache. The database is comfortable up to 500 QPS. (a) What hit rate do you need? (b) Ops reports "hit rate dipped from 98% to 95% overnight — probably fine?" Is it? (c) What dashboard line would have made this a non-event?
(a) Misses must stay under 500 of 20,000: you need ≥ 97.5%. (b) Not fine — an incident in slow motion. At 98%, misses = 400 QPS, near the line; at 95%, misses = 1,000 QPS, double the comfort zone. A "3-point dip" was a 2.5× jump in database load; the percentage framing hides it. (c) Absolute miss QPS with an alert near 400 — the miss line is in the same units as the database's capacity.
What breaks: a nightly job at 3:00am pre-warms all 5 million user profiles, TTL 6 hours, no jitter. The app sends its daily push notification at 8:55am. Walk the timeline to the incident, then fix it two ways.
3:00am — 5M keys filled in minutes; expiries all land ~9:00am. 8:55am — push fires, peak begins, cache still warm. 9:00–9:05am — nearly every key expires during the day's busiest minutes; hit rate cliffs; the database takes peak traffic uncached. Cruelest part: it looks like the push's fault, but the cause fired six hours earlier. Fix one: jitter — ttl = 6h + random(0, 1h). Fix two: stale-while-revalidate, so expiry means "refresh in background," never "everyone misses at once." (Bonus: question whether pre-warming 5M is needed — cache-aside warms the actives organically.)
Commit and justify: the PM wants view counts on videos, ~100,000 increments per second at peak. One teammate proposes writing each view to the database ("source of truth"); another proposes write-back through Redis. Pick one, name the loss scenario, and defend it.
Write-back, committed: INCR video:{id}:views in Redis (atomic, memory-speed), flush aggregated deltas every few seconds. Per-view DB writes at 100K/s means heavy infrastructure to perfect a number nobody needs perfect. Loss scenario: a Redis node dies between flushes and a few seconds of increments vanish — invisible, harmless, because a view count carries no promise. The moment a promise appears — a paid view, a "watched" marker — write-back is forbidden: durable storage before acknowledgment. Write-back is a durability decision disguised as a performance pattern.
Spot the bug: server A reads user 42 from the DB (gets v1), then stalls on a GC pause. Server B updates user 42 to v2 and deletes user:42 from the cache. Server A wakes and runs cache.set("user:42", v1). What state results, for how long, and what are two real fixes?
The cache holds v1 — stale — with a fresh full TTL, while the database holds v2; every reader sees the old profile until the TTL expires. This is the "stale set" race: delete-on-write happened and did not prevent it. Fixes: (1) leases / conditional fill — the cache issues a fill token on miss, B's delete voids A's token, A's late set is refused (Facebook's solution; compare-and-set on a version field also works); (2) modest TTLs even on invalidated keys — not prevention, but a ceiling on the damage. (Versioned keys sidestep it entirely — A writes to a key nobody reads anymore.)
Your photo app runs on servers in Virginia. It's fast — you've profiled it. The server builds a response in 40 milliseconds. Then a friend in Delhi messages you: "Nice app, but why does every photo take three seconds to load?"
You check everything. Database: fine. Cache hit rate: great. Server CPU: bored. The app is fast — in Virginia. Your friend's request is crossing twelve thousand kilometers of ocean floor, twice, per round trip — and a fresh HTTPS connection needs three or four round trips before the first image byte even starts moving.
No amount of code optimization fixes this. You are not fighting your stack. You are fighting the speed of light — and light is not going to get faster. This chapter is about the only move that works: stop making the bytes travel. Put copies of them near every user (that's a CDN), and put the originals somewhere built to hold trillions of files forever (that's blob storage). Almost every real system you'll design in Part 3 — Instagram, YouTube, Dropbox — stands on these two legs.
Light in a fiber-optic cable travels at about 200,000 km per second — two-thirds of its speed in a vacuum, because glass slows it down. That sounds fast until you do the division. Virginia to Delhi is roughly 12,000 km. One round trip: 24,000 km ÷ 200,000 km/s = 120 milliseconds — and that's the physics-textbook best case. Real cables don't run in straight lines, and every router along the path adds a little, so the real round-trip time (RTT) is more like 180–250 ms.
Now count round trips. A brand-new HTTPS connection costs: one round trip for the TCP handshake (the "hello, are you there?" that starts every connection), one more for the TLS handshake (agreeing on encryption), and one for the actual request and response. Three round trips minimum — call it 600 ms from Delhi before the first image byte arrives. A page with images on a few separate connections stacks these costs. That's your three seconds, and your server's 40 ms was never the story.
| Distance | Typical RTT | Feels like |
|---|---|---|
| Same city | 1–5 ms | Instant |
| Same continent (coast to coast) | 60–80 ms | Snappy |
| Across an ocean | 150–250 ms | Noticeable pause |
| Antipodes (Sydney ↔ Virginia) | 200–300 ms | Painful, and it multiplies per round trip |
The senior takeaway from this table: latency across distance is not an engineering problem, it's geography. The only fix is to shorten the trip. So the industry built exactly that.
A CDN — content delivery network — is a company (Cloudflare, Akamai, Fastly, CloudFront) that runs hundreds of small data centers around the world, called edge PoPs (points of presence). Delhi has several. So does Sydney, São Paulo, Lagos. Each PoP is a big shared cache, sitting 5–20 ms from the users near it.
You met caching in the last chapter. A CDN is that same idea with a passport: it caches your responses geographically, close to whoever is asking. The mechanics:
photos.yourapp.com/cat.jpg. DNS steers that hostname to the nearest PoP instead of your server (CDNs do this with anycast or geo-aware DNS — you saw how that routing works in chapter 1.1).cat.jpg. It replies in ~20 ms. Your Virginia server never hears about it.A CDN is the corner shop, and your origin is the dairy farm. Nobody drives 300 km to the farm for milk — the shop stocks it locally, and when the fridge runs low, the shop restocks from the farm. The TTL is the expiry date on the carton: the shop will confidently sell it until that date, then check with the farm for a fresh one. Where the analogy breaks: a shop pays for stock up front, while a pull CDN only fetches after someone asks — the first customer in each neighborhood waits for the delivery truck.
How long does the edge keep a copy? You tell it, with a TTL (time to live), usually via the Cache-Control response header. max-age=3600 means "any cache may serve this for an hour without asking me again." The trade-off is the same one from the caching chapter: long TTL = fewer origin hits but staler content; short TTL = fresher but your origin works harder.
What counts as "the same file"? The cache key — by default the URL, plus any headers or query parameters you choose. Get it wrong in either direction and things go sideways. Include too much (say, a per-user session cookie) and every user misses — hit rate quietly hits 0% while your origin takes full traffic. Include too little and you serve the wrong bytes: if ?size=small isn't in the key, the desktop user gets the thumbnail. The truly dangerous version is caching a personalized response under a shared key — then user A is served user B's page.
On December 25, 2015, Steam users started seeing other people's account pages — strangers' email addresses, purchase histories, partial payment info. Valve took the store offline. It wasn't a hack: Steam was under a denial-of-service attack, and a caching configuration deployed to absorb the traffic was wrong. For about an hour, personalized store pages were cached and served to other users as if they were shared static files; Valve's statement put it at around 34,000 users exposed. The lesson is precise: a cache key is a security boundary. Anything personalized must either never be cached at the edge or be keyed so it can only ever reach its owner. One wrong caching rule turned a traffic problem into a privacy incident on the store's busiest day of the year.
Now the harder question, the one interviewers love: you pushed a broken image, and it's cached at 300 PoPs with a 24-hour TTL. How do you take it back? Two tools:
app.3f9a2c.css — and serve it with Cache-Control: max-age=31536000, immutable (a year, never revalidate). When content changes, the URL changes, and every cache on Earth fetches the new file naturally. You never invalidate because you never reuse a URL. Every modern web build pipeline does exactly this.The rule of thumb: version what you can, purge what you must (a leaked file, a legal takedown, a cached error page).
Everything above describes a pull CDN: lazy, fills itself on demand, the default for nearly everyone. Its weakness is the cold cache — a new file, a just-purged file, or a new PoP means a burst of misses hitting your origin at once (good CDNs soften this by collapsing concurrent misses for the same key into one origin fetch).
A push CDN flips the direction: you upload to the edges ahead of time, before anyone asks. More work — you decide what goes where and pay for edge storage — so it only pays off when demand is predictable and the cost of a miss is brutal. There is one company for whom both are extremely true.
By 2011, Netflix streaming was on its way to being a double-digit share of all downstream internet traffic (Sandvine has measured it around 15% globally in recent years) — too big to keep renting commercial CDNs. So in 2012 Netflix built its own, Open Connect, with a twist: it ships Open Connect Appliances — servers stuffed with hundreds of terabytes of storage — to internet service providers, free, to install inside their own networks. Your stream often never touches the public internet; it comes from a box in your ISP's building. The second twist is the push model: every night, during off-peak hours, each appliance downloads the titles Netflix predicts that region will watch tomorrow, so the evening peak is served from disks a few kilometers away. ISPs win too — Netflix bytes stop clogging their expensive upstream links. Today effectively all Netflix video is served from Open Connect appliances in thousands of locations across 100+ countries. That's why a global simultaneous release doesn't melt the internet: the show was already everywhere before anyone pressed play.
Interviewers ask: "my API responses are personalized — is the CDN useless?" No. Dynamic acceleration: the user still connects to the edge 20 ms away, so the TCP and TLS handshakes happen over 20 ms instead of 220. From there the request rides the CDN's private backbone to your origin over connections that are already open and warm. Zero caching, and you've still turned three slow round trips into three fast ones plus one long one.
A CDN edge serves anyone who asks — that's the point. But what about paid course videos, or private photos? You can't just put them on public URLs and hope nobody shares the link.
The answer is a signed URL: your app server combines the URL, an expiry time, and a secret key into a cryptographic signature appended as a query parameter. The edge verifies signature and expiry locally — no call to your origin — and serves or rejects. Your app decides who gets a link (only logged-in, paid-up users); the CDN enforces that the link is genuine and fresh, and a leaked link dies at expiry. For content made of many files — a video split into hundreds of segments — signing each URL is clumsy, so CDNs offer signed cookies: one signed grant authorizing everything under a path like /course-42/*.
The CDN caches copies. Copies of what, stored where? Your first instinct — the app server's disk — dies the moment you have two app servers (which one has the file?) or one server failure (where did the photos go?). Your second instinct might be the database. Hold that thought — we're about to kill it properly.
What you want is a blob store (blob = binary large object: images, videos, PDFs, backups — any opaque pile of bytes). The model the whole industry copied is Amazon S3, launched 2006: you create a bucket (a namespace) and put objects into it under keys — strings like photos/2026/08/cat.jpg. The slashes are cosmetic: the namespace is flat, keys just look like folder paths. The API is deliberately tiny — PUT, GET, DELETE, LIST — and there's no "edit byte range 100–200"; you replace an object whole. In exchange you get the numbers that matter: objects up to 5 TB, effectively infinite capacity, pay-per-GB, and eleven nines of durability (99.999999999% — redundant copies across multiple data centers make losing an object a once-in-millions-of-years design goal). Availability is a different, lower number, and a senior keeps the two words apart: durability is "is my data still alive somewhere," availability is "can I reach it right now."
One fact to state correctly, because it changed: since December 2020, S3 is strongly consistent — after a successful write, any read anywhere immediately sees the new object, including overwrites and LISTs. For S3's first fourteen years, overwrites were only eventually consistent (a read just after a write could return the old version), and a generation of interview prep material still says so. Saying "S3 is eventually consistent" in 2026 is a small but real credibility ding.
So why not blobs in the database? Because database storage is the most expensive kind you own. Every gigabyte of photo bytes in Postgres sits in the engine's carefully managed pages, gets copied to every replica, bloats every backup, and shoves hot rows out of the buffer pool (the RAM cache that makes queries fast). And for all that cost, the database gives blobs nothing back — you can't index, query, or join on the middle of a JPEG. You'd be paying for a scalpel and using it as a shovel.
The pattern every serious system uses instead:
photos table (Postgres) blob store (S3)
-------------------------- -----------------------
id: 8814 key: photos/8814/orig.jpg
owner_id: 442 (the actual 3 MB of bytes)
s3_key: "photos/8814/orig.jpg"
width, height, created_at, ...
The database holds metadata — who owns it, when it was made, where it lives — everything you query and join on. The store holds the bytes; the row points at the object by key. This split shows up in every Part 3 design that touches media, and interviewers treat "blobs in the DB" as a genuine red flag.
It's a coat check. The cloakroom (blob store) holds the heavy coats; you hold a small numbered ticket (the database row with the key). Nobody rummages through a rack of coats to find yours — they look up the ticket number and go straight to the hook. And you'd never staple the coat to the ticket: the ticket stays light precisely so the front desk can flip through thousands of them per second. The analogy is honest about one more thing: lose the ticket and the coat is orphaned — which is why the metadata row and the object need to be created and deleted together, carefully.
On February 28, 2017, an Amazon engineer debugging a slowdown in S3's billing system in us-east-1 — AWS's oldest, busiest region — ran a playbook command to take a few servers out of service. One parameter was mistyped, and the command removed a large chunk of capacity instead, including servers running S3's index subsystem (which knows where every object lives) and its placement subsystem (which decides where new objects go). Both needed full restarts. Neither had been fully restarted in years, and S3 had grown enormously in that time; recovery took roughly four hours. The outage cascaded through everything built on S3 — Trello, Quora, IFTTT, Slack file uploads, countless sites whose images simply vanished. The quietly damning detail: AWS's own status dashboard couldn't show its red warning icon, because the icon was stored in S3. The fixes were pure senior-engineering hygiene: the capacity-removal tool now refuses to drop below a safety floor, and the subsystems were partitioned into smaller cells so no restart is ever that big again. Two lessons for interviews: ops tooling needs guardrails exactly like APIs do, and when half the internet keeps its bytes in one system in one region, that system's blast radius is half the internet.
Downloads sorted. Now uploads, where there's a subtler trap. The obvious flow — client uploads to your app server, which writes to S3 — makes your servers a pipe. Run the numbers: 2 million uploads a day at 3 MB each is 6 TB flowing through your servers daily, ~550 Mbps average and several times that at peak, each in-flight upload holding a connection and buffer memory for seconds. Your fleet scales with upload bandwidth while adding zero value to the bytes — it receives them and immediately forwards them.
The fix is the presigned upload URL — the CDN section's signature trick, pointed the other way. Your server uses its S3 credentials to sign a short-lived permission slip: "the bearer may PUT one object to exactly this key, with this content type, within 10 minutes." The client then uploads directly to S3. Your server handles a tiny JSON exchange; the 3 MB rides S3's bandwidth, not yours.
Two operational details separate a working answer from a hand-wave. First, constrain the slip: pin the exact key, content type, and a size limit in the signature — otherwise you've handed the client a blank check to fill your bucket. Second, don't trust "done." Clients crash, lie, and lose connectivity. Create the metadata row only after S3 confirms the object exists — via a HEAD check when the client reports back or, better, S3's event notification firing on object creation.
A 4 GB video over one HTTP PUT is a bad bet: S3 caps single PUTs at 5 GB anyway, and one dropped packet at 99% means restarting from zero — on a phone connection the upload may never finish. Multipart upload fixes this: initiate a session, split the file into parts (5 MB–5 GB each, up to 10,000), upload parts independently — in parallel for speed, retrying only what failed — then send a "complete" call and S3 assembles the object. A flaky connection now costs one 10 MB part, not 4 GB of progress.
The habit that marks people who've run this: abandoned uploads (user quits at part 3 of 400) silently keep their parts, and you pay for invisible half-files forever. S3 has a lifecycle rule to auto-abort incomplete multipart uploads after N days — set it on day one. Which brings us to lifecycle rules in general, because storage has a temperature.
Data ages predictably. A photo gets most of its views in week one; by month six it's opened once in a blue moon; after two years it exists mostly as an archive. Paying hot-storage prices for frozen data is burning money, so S3 (and every peer) offers storage classes — same API, different price/retrieval trade-off:
| Class | ~Price per GB-month | Retrieval | Use for |
|---|---|---|---|
| Standard (hot) | ~2.3¢ | Instant, free | Actively served content |
| Infrequent Access (IA) | ~1.25¢ | Instant, per-GB fee | Older content, still occasionally read |
| Glacier tiers (archive) | ~0.4¢ down to ~0.1¢ | Instant to hours, fees; minimum storage 90–180 days | Compliance, backups, "probably never" |
A lifecycle rule automates the demotion: "after 30 days, move to IA; after 365, to Glacier; delete logs after 7 years." At petabyte scale this is real money — 1 PB in Standard is about $23,000/month; in Glacier Deep Archive, about $1,000.
The gotcha interviewers enjoy: cold classes charge retrieval fees and enforce minimum storage durations (30 days in IA, up to 180 in Deep Archive). Move data that's still being read into IA and the retrieval fees can exceed your savings. Tiering decisions need measured read rates per age cohort — access numbers, not vibes.
One pattern assembles this whole chapter, and it appears in Instagram, YouTube, WhatsApp, and half of Part 3. A user uploads a photo or video; you need it resized into variants (thumbnail, feed, full size) or transcoded into multiple video qualities — which for video takes minutes of CPU, not milliseconds. None of that can happen on the request path; the user won't hold the line while you transcode 4K.
Walk the path: the client uploads the original via presigned URL. S3's object-created event drops a message onto a queue (a buffer that holds work until a worker is free — chapter 1.11 dissects these). Workers pull jobs, produce the variants, write them back to S3 under versioned keys, and mark the metadata row "ready." The CDN serves the variants with year-long immutable TTLs. Meanwhile the app shows the uploader "processing" — or, the common product trick, a locally rendered preview so it feels instant.
Why a queue instead of resizing synchronously after upload? Because upload traffic spikes (everyone posts at 9 pm) while transcoding capacity is expensive and finite. The queue absorbs the spike; workers chew through it at their own pace; if a worker dies mid-job, the message returns to the queue and another worker retries. For images there's a legitimate alternative — resize on the fly at request time and let the CDN cache the result, trading first-request latency for zero wasted preprocessing. For video there is no debate: transcoding is minutes of work and always goes through the pipeline.
CDN and blob storage rarely headline a question — they're assumed vocabulary inside every media-flavored design, and seniors are graded on deploying them unprompted with correct mechanics: presigned upload and CDN-fronted download drawn without being asked. Expect probes like:
A junior asks: "Why does everyone store images in S3 and serve them through a CDN, instead of just keeping them in our database and serving them from our API?" Answer in five or six sentences.
Two different problems, two specialized tools. A database is built for small structured rows you query and join; a 3 MB photo is an opaque blob it can do nothing clever with, yet it would bloat every backup, every replica, and the RAM cache that keeps real queries fast. So the photo goes to S3 — unlimited blobs, cheap, almost never loses one — and the database keeps a small row pointing at it. Serving is a physics problem: our servers live in one place, and a user on another continent pays 200 ms of light-travel time per round trip no matter how fast our code is. A CDN keeps copies in hundreds of cities, so nearly every user downloads from a cache a few milliseconds away, and our servers only ever see the occasional miss. Database for facts, S3 for bytes, CDN for distance.
Your app's users upload 2 million photos/day, averaging 3 MB. (a) If uploads are proxied through your app servers, what average and peak bandwidth do the servers carry? (b) How much new storage per year, and roughly what does it cost per month in S3 Standard by year's end? Assume peak is 3× average.
(a) 2M × 3 MB = 6 TB/day ÷ 86,400 s ≈ 70 MB/s ≈ 560 Mbps average, ~1.7 Gbps at peak — through servers that add nothing to those bytes. That number is the argument for presigned URLs. (b) 6 TB/day × 365 ≈ 2.2 PB/year of originals. At ~2.3¢/GB-month that's ≈ $50K/month by year's end and climbing — which is why the lifecycle-to-IA conversation isn't optional at this scale. Both answers changed a design decision; that's what made them worth computing.
Your CDN serves 100,000 requests/sec at a 99% hit rate, and your origin comfortably handles about 3,000 requests/sec. A config mistake drops the hit rate to 94%. What happens, and what does this teach about how to monitor a CDN?
At 99%, origin sees 1% of 100K = 1,000 req/s — comfortable. At 94%, 6,000 req/s: six times the load, double what the origin survives. A 5-point dip 6x'd origin traffic and took you down, because the CDN was silently absorbing 99% of true demand. Lessons: alert on hit rate and origin QPS directly (QPS moves before origin CPU does), and capacity-plan the origin for a realistic worst-case hit rate, not the sunny-day one. Same amplification is why a mass purge during peak hours is a self-inflicted outage.
Design the upload flow for a Dropbox-like app where users upload files up to 50 GB from laptops on unreliable Wi-Fi. Name the mechanisms and the failure handling.
Server authorizes and initiates a multipart upload, issuing presigned URLs per part (say 64–128 MB each — 50 GB is far over the 5 GB single-PUT cap anyway). Client uploads parts in parallel; on connection loss it asks which parts completed and resumes, re-uploading only failures. After the final part, a complete call assembles the object, and an object-created event — not the client's word — triggers metadata finalization. Round it out with part checksums, presigned URLs scoped to exact part and expiry, and a lifecycle rule auto-aborting incomplete uploads after ~7 days so abandoned halves don't bill you forever.
A chat app stores 1 PB of media. Analytics: read heavily for 30 days after upload, rarely from day 30–365 (~1% re-read per month), almost never after a year. Design the lifecycle policy and sanity-check it against the retrieval-fee gotcha.
Standard for days 0–30 (hot reads need free, instant retrieval). At day 30, transition to IA: storage drops from ~2.3¢ to ~1.25¢/GB-month, and the 1%/month re-read rate at ~1¢/GB retrieval adds a negligible fee against the savings on the other 99%. After 365 days: Glacier Instant (~0.4¢) if old media must open instantly in chat history; Deep Archive only if the product tolerates an hours-long "retrieving your file" state. The check that matters: the IA move wins because the measured re-read rate is 1% — at 50% re-reads, retrieval fees would eat the savings. Tiering decisions come from access numbers, never from data age alone.
It's 2:47am. Your phone lights up: db-prod-1 unreachable. You ssh in. Nothing. The machine that holds every user, every order, every password hash your company has ever stored is not answering, because its disk controller died somewhere between 2:41 and 2:43. There is exactly one copy of your data, and it is inside a metal box that no longer turns on.
Maybe the disk is recoverable. Maybe your nightly backup restores cleanly — losing only everything since midnight. Maybe. You'll find out over the next six hours, in front of a status page turning red, doing arithmetic on how many signups happened since the last backup.
Here's the math that makes this inevitable rather than unlucky. A server disk fails at roughly 1–2% per year — the annualized rate cloud storage companies like Backblaze publish from fleets of hundreds of thousands of drives. One disk, ignorable. Two hundred disks: expect two to four funerals a year. Ten thousand: disks die weekly, on schedule, forever. And disks are just the start — kernels panic, power supplies pop, whole data-center zones go dark. At any real scale, "a machine died" is not an incident. It's Tuesday.
The only real defense is embarrassingly simple: keep more than one copy. That's replication — the same data, maintained on several machines, kept in sync automatically. Simple to say. The rest of this chapter is about why keeping copies in sync is one of the deepest problems in distributed systems, and how it bites users, engineers, and — in one famous case — all of GitHub.
The most common design for keeping copies in sync, behind almost every Postgres and MySQL deployment you'll ever touch, is leader-follower replication (also primary/replica, or the older master/slave). One machine is the leader: it accepts every write. The others are followers: full copies that typically serve reads but never accept a write directly from a client.
How do followers stay in sync? Every time the leader commits a change, it appends a record of that change to an ordered, append-only file — Postgres calls it the write-ahead log (WAL), MySQL calls it the binlog. Followers hold an open connection to the leader, stream that log, and replay each entry in the exact order the leader wrote it. Same starting state, same changes, same order — each follower converges to a byte-for-byte copy. This is log shipping, and the log-of-ordered-changes idea will follow you through this whole book (it's also the soul of Kafka, Chapter 3.17).
Why ship a log of effects instead of forwarding the SQL? Because SQL isn't deterministic: NOW() produces a different timestamp on every machine that runs it. The log records what actually changed — "row 4182, new value 02:47:03" — so every copy lands on identical bytes.
The leader is the accountant who keeps the firm's official ledger. Two clerks sit nearby, each copying every new line into their own ledger, in order, as it's written. Anyone can answer "what's the balance?" from any clerk's book — but new entries go into the official ledger first, always. If the accountant's book burns, a clerk's copy saves the firm. The catch, and it's the catch this whole chapter turns on: a clerk is always a few lines behind the pen.
One thing this does not give you: write scaling. Every write still funnels through one leader, and — read this twice — every follower must also replay every write to stay a copy. Three replicas means the same write performed three times, not shared three ways. Replication buys survival and read capacity. Writes scale by splitting data, not copying it — sharding, Chapter 1.9.
Look at the diagram again. The leader has written locally and is about to tell the client "saved." Does it wait for a follower to confirm first? That one decision defines what your database promises.
Synchronous replication: the leader waits until a follower confirms the write before acknowledging the client. The data provably exists on two machines before anyone was told it's safe — a leader death loses nothing acknowledged. The price: every commit pays a round trip to the follower — under a millisecond inside one data center, but 60–70ms between the US coasts and 150ms-plus to Asia. Worse: if the synchronous follower dies or hiccups, the leader must stop accepting writes. Fully synchronous to all followers means any one machine's bad day halts your database — which is why essentially nobody runs that.
Asynchronous replication: the leader acknowledges immediately and ships the log in the background. Fast, smooth, the default almost everywhere. The price is a small, sharp lie: there is always a tail of writes — milliseconds' worth, sometimes seconds' — that the leader has acknowledged but no follower has yet. If the leader dies right then, that tail is gone. You told the user "saved." It wasn't. For a like, who cares. For a payment, careers end.
Semi-synchronous replication: the grown-up compromise. The leader waits for at least one follower to confirm it has received the log entry (not necessarily applied it yet), then acks. Every acknowledged write exists on two machines, and one slow follower doesn't halt writes while another can confirm. MySQL's semi-sync mode and Postgres's synchronous_standby_names = ANY 1 (...) both do this. Keep the confirming follower in the same region and the cost is a millisecond or two per commit — a bargain for "we don't lose acknowledged data when one machine dies."
You're sending a signed contract. Async is dropping it in a postbox and immediately telling your boss "sent and done" — almost always fine, until the one day the mail van burns. Sync is a courier who won't let you say "done" until the recipient has signed for it — safe, slow, and if the courier is stuck in traffic, your whole day stops. Semi-sync is getting confirmation that the post office has it: not yet delivered, but no longer in only one place. That's exactly what semi-sync confirms — receipt on a second machine, not application.
| Mode | Commit latency cost | Leader dies — what's lost? | When a follower dies |
|---|---|---|---|
| Async | None | The unshipped tail of acknowledged writes — seconds of data, gone | Nothing; leader doesn't wait anyway |
| Semi-sync (ANY 1) | ~1 round trip to nearest follower (<2ms same region) | Nothing acknowledged | Fine while any confirming follower survives |
| Fully sync | Round trip to slowest follower, every commit | Nothing acknowledged | Writes freeze until it recovers |
The senior framing: this is a durability-versus-latency budget, spendable differently per system. A committed answer: semi-sync to one follower in the same region (survives machine death, ~1ms cost), async to another region (disaster recovery without a 70ms cross-country toll on every commit). Say that with the numbers attached and you sound like someone who has carried the pager.
Async and semi-sync share a property: followers run slightly behind the leader. That gap is replication lag. Healthy, it's tens of milliseconds. Under a traffic spike, a bulk migration, a long transaction, or a follower with slow disks, it stretches to seconds or minutes — and many databases replay the log on followers with less parallelism than the leader used to generate it, so a leader shrugging off write load can be feeding a follower that simply cannot keep up.
Lag would be invisible internals except that you're reading from those followers. Here's the classic bug, the one every engineer eventually ships. A user posts a comment; the write goes to the leader. The page reloads, and the read gets load-balanced to a replica 200ms behind — which doesn't have the comment yet. The user watches their comment appear, then vanish on refresh. They didn't see "replication lag." They saw your app eat their words, and filed a one-star review about it.
Notice something important: the system isn't broken. Every component is doing its job. The bug lives in the gap between copies — which is why you can't fix it with a patch, only with a policy. Three policies matter:
Read-your-writes. The guarantee: a user always sees their own writes. Cheapest implementation: after a user writes, route that user's reads to the leader for a short window (say, 10 seconds) — a sticky "recently wrote" flag in their session. More precise: every position in the replication log has an identifier (Postgres calls it an LSN — log sequence number). Record the LSN of the user's write in their session; a replica serves their reads only if it has replayed past that LSN, else the read waits a beat or falls through to the leader. Precise, but now your app is lag-aware plumbing. Start with the pin-to-leader window; graduate to LSN tracking when the leader feels the pinned load.
Monotonic reads. A sneakier bug: a user who wrote nothing refreshes twice, the first read hits a fresh replica, the second a lagged one — and someone else's comment they just read disappears. Time ran backwards. Monotonic reads forbids this: a user may see stale data, but never older data than they've already seen. Standard fix: session stickiness — hash each session to one replica, so one user rides one timeline. (Volunteer the caveat before the interviewer does: when that replica dies, its users hop timelines, so stickiness softens the bug rather than erasing it.)
Honest staleness. For plenty of surfaces — view counts, dashboards, search results — the senior move is to serve stale data on purpose and say so. Freshness is a budget. Spend it where users notice: their own writes.
Replication lag doesn't just confuse users — it pages tired humans, and tired humans touch primaries. On 31 January 2017, a spam wave hammered GitLab.com's Postgres and replication to the secondary fell so far behind it broke entirely. Late that night, an exhausted engineer rebuilding the secondary meant to wipe its stale data directory — and ran rm -rf on the primary. Roughly 300GB, gone. Then the safety nets failed roll call: the automated pg_dump backups had been silently producing nothing for weeks (a Postgres version mismatch), the S3 backups were empty, and disk snapshots weren't enabled for the database servers. What saved them was a manual LVM snapshot taken about six hours earlier for an unrelated staging refresh. GitLab restored from it — live-streaming the recovery on YouTube — and permanently lost about six hours of issues, merge requests, comments, and new accounts (Git repository data lived elsewhere and survived). Two tattoo-grade lessons: lag is an operational hazard, not just a UX one; and replication is not backup — a replica faithfully replays your deletes. Copies protect you from dead hardware. Only backups protect you from commands.
Replication's promise is that when the leader dies, a follower takes over. That handover is failover: notice the leader is dead, promote a follower, repoint everything at it. Three steps that sound trivial and are anything but.
Detection is the first trap. Machines don't announce their death; they go quiet. So we use heartbeats: no answer for N seconds, declare it dead. But "quiet" has many causes — a garbage-collection pause, a saturated link, a brief partition — and from the outside, a dead leader and a slow leader look identical. Set N too small and ordinary hiccups trigger failovers. Set it too large and every real death means minutes of downtime. There is no correct N; only a trade-off you must own.
Promotion is the second trap. You promote the most caught-up follower — but under async replication, even it may be missing the dead leader's final seconds of acknowledged writes, which get discarded to keep histories consistent. Acknowledged, then discarded: the async contract collecting its fee at the worst moment.
Then there's the third trap, the one with a horror-movie name. Suppose the old leader wasn't dead — just unreachable for a while. Your automation promotes a follower. Now two machines both believe they're the leader, and clients on different sides of the hiccup are writing to both. This is split brain: two divergent histories of the truth, growing apart with every write, with no automatic way to merge them. Balances updated on both sides; the same auto-increment ID handed to two different rows. Cleanup is manual archaeology, row by row.
A company's head office goes silent for one minute — phones down. Regional office B, following the emergency handbook, declares its manager acting CEO. But head office was never gone, just unreachable, and its CEO kept signing deals the whole time. Now two CEOs have each signed a minute of contracts, some for the same clients, and no lawyer on earth can automatically merge them. The fix in real companies and real databases is the same: before anyone takes the crown, make certain the old ruler can't act — revoke their signing authority, not just their phone line.
That "make certain" step has a name: fencing — forcibly preventing the old leader from doing damage before the new one takes over. Blunt version: STONITH, "shoot the other node in the head" — cut the old machine's power via its management interface. Elegant version: every leadership term gets an increasing epoch number, and storage and clients reject writes stamped with a stale epoch, so a zombie leader's writes bounce off. Chapter 2.3 builds fencing properly on consensus. For now, burn in the rule: promotion without fencing is a split brain scheduled for later.
On 21 October 2018 at 22:52 UTC, routine maintenance to replace failing optical equipment briefly severed the link between GitHub's US East Coast data center and its East Coast network hub. Connectivity returned in 43 seconds. That was enough. Orchestrator — GitHub's automatic MySQL failover system — saw East Coast primaries unreachable and did exactly what it was built to do: promoted replicas in the West Coast data center and redirected writes there. But cross-country replication was asynchronous, so the East primaries held a tail of acknowledged writes that had never reached the West — and the newly promoted West primaries immediately began accepting fresh writes of their own. Forty-three seconds of network trouble had produced two databases on opposite coasts, each containing writes the other lacked. GitHub refused to guess at a merge: they kept the site degraded, paused webhook delivery and Pages builds (millions of deliveries queued up), restored the East Coast databases from backup, replayed replication until the coasts converged, and reconciled the orphaned East Coast writes. Total time degraded: 24 hours and 11 minutes — from 43 seconds of partition. The automation didn't malfunction; it did precisely what it was configured to do, faster than any human could evaluate whether it should. That's the lesson seniors carry: automatic failover fails in exactly the scenarios hardest to test — partial, asymmetric, brief. GitHub's remediations pushed toward tighter fencing and topology rules so a short blip could never again trigger a cross-country promotion.
So should failover be automatic? Commit to a position: automate detection, preparation, and everything reversible; be extremely careful automating promotion itself. Many strong teams have automation detect, pick the candidate, pre-stage the runbook — and require a human to press the button, because promotion is irreversible and "leader unreachable" is ambiguous. Teams that fully automate — at large fleet sizes you must; humans don't scale to a failover a day — pair it with real fencing and a consensus-based view of who's alive: a majority of independent observers must agree the leader is gone before anyone is crowned. Majorities and quorums get their full treatment in Chapter 1.10.
Beyond survival, followers earn their keep serving reads — and most systems are read-heavy, often 10:1 or 100:1. Point feed reads, profile views, and analytics at replicas; reserve the leader for writes and the reads that truly cannot tolerate staleness. Five replicas, roughly five times the read throughput — the standard second act of every scaling story, right after "add a cache."
But replication factor is not infinite read scale, and interviewers love candidates who know where the ceiling is. Three limits:
And the ceiling on writes is absolute: no number of replicas adds a single write per second, because every write still flows through one leader and then into every copy. When the leader's write throughput is the wall, you stop copying data and start splitting it — sharding, next chapter.
Multi-leader replication puts a leader in each region: users write locally at local latency (Sydney users no longer pay 200ms to a Virginia leader), and leaders exchange changes asynchronously. The bill arrives immediately: two users can now edit the same row in two regions in the same second, and both writes are accepted. Someone must decide who wins — last-write-wins (quietly discards data), custom merge logic, or punting to the application. Conflict resolution is genuinely hard and unavoidable; it gets full treatment in Chapter 2.5. The one-line rule: don't take multi-leader for durability — take it only when multi-region write latency forces you to, with a conflict story you can defend.
Leaderless replication — the Dynamo style, used by Cassandra and Riak — abolishes the leader entirely. The client writes to several replicas at once, done when W of N confirm; reads query R replicas and take the newest version. Choose W + R > N and every read overlaps at least one replica that saw the latest write — no leader, no promotion, no failover drama, at the price of a subtler consistency story. That arithmetic — quorums — anchors Chapter 1.10.
Replication is rarely its own question — it's the trapdoor under every design you draw. Put a database box on the board and assume these are coming:
A junior asks: "We pay for five database servers but one does all the writing, and yesterday my test comment vanished when I refreshed. Why?" Explain both, in five or six sentences.
One machine — the leader — takes every write and streams a log of changes to the other four, which replay it to stay identical copies. We do this so a dead machine costs us nothing (a copy takes over) and so the four followers can serve our read traffic, which is most of our traffic. But the copies run slightly behind the leader — replication lag, usually milliseconds, sometimes seconds under load. Your comment went to the leader; your refresh read from a follower that hadn't replayed it yet, so it looked deleted, then reappeared once the follower caught up. The standard fix is read-your-writes: for a little while after you write, we send your reads to the leader, so you always see your own changes. And one warning while we're here: those five copies are not backups — a bad delete replicates to all five in milliseconds.
Your leader is in Virginia. Compliance wants synchronous replication to Oregon (~70ms round trip) so no acknowledged write can ever be lost to a regional disaster. Checkout commits 5 sequential writes. What happens to checkout latency, and what do you counter-propose?
Each commit waits a 70ms cross-country confirmation: 5 × 70 = ~350ms added to checkout. And when Oregon has a bad network day, Virginia stops accepting writes: you've coupled uptime to two regions' health. Counter-proposal: semi-sync to a follower in a different availability zone of the same region (~1–2ms per commit) covering machine and AZ death, plus async to Oregon for regional disaster — accepting that a true region-killing event may lose the last few seconds. Then make the trade explicit: "zero loss even if Virginia is destroyed mid-transaction" costs 350ms per checkout; "zero loss unless an entire region dies, then seconds at most" costs 10ms. Numbers turn a policy fight into a decision.
Users report: "I post a comment, refresh, it's gone, refresh again, it's back." Sketch two fixes, and for each name the cost and the failure mode that remains.
Fix 1 — pin-to-leader window: after any write, route that user's reads to the leader for ~10s via a session flag. Cost: extra leader read load, scaling with write rate. Remaining hole: the window is a guess — lag longer than 10s (migrations, incidents) resurfaces the bug. Fix 2 — LSN tracking: store the log position of the user's last write in their session; replicas serve their reads only if replayed past it, else fall through to the leader. Cost: lag-aware plumbing in the read path. Remaining hole: cross-device — comment on your phone, check on your laptop, and the laptop session carries no LSN; closing it needs server-side per-user write positions. Interview bonus: add session stickiness so each user rides one replica's timeline — monotonic reads, which also kills the "someone else's comment flickers" variant.
A teammate ships a cron: "if the leader misses heartbeats for 10 seconds, auto-promote the freshest follower and update DNS." List what can go wrong, in order of how badly you'll be paged.
(1) Split brain, the career-limiting one: a 15-second GC pause or partition isn't death — the script promotes while the old leader is alive, and with no fencing both take writes; you're now GitHub in October 2018. (2) Acknowledged-write loss: the "freshest follower" under async can still be seconds behind; promotion silently discards the tail, and DNS caching keeps some clients writing to the zombie even longer. (3) Flap storms: a wobbly network trips repeated promotions. (4) Cold-start stampede: the promoted follower inherits full traffic with cold caches at the worst moment. Minimum fixes: fence before promote (cut power or invalidate the epoch), require agreement from multiple observers across failure domains before declaring death, and repoint via a connection proxy you control, not DNS TTLs you don't.
You run 1 leader + 2 replicas at 80% read capacity; the workload is 40% writes and read traffic is about to double. The intern proposes "add 4 more replicas." Does it work? Do the reasoning.
Mostly no — and the reason is the exercise. At 40% writes, a big slice of every replica's capacity goes to replaying the leader's write stream before it serves one read: each added replica brings maybe 60% of a machine's worth of usable read capacity. Four more replicas roughly doubles usable read capacity — barely covering the doubling, zero headroom — while adding four more lagging timelines, four more log streams from the leader, and four more machines to operate. Better order: profile the reads first. Hot-key traffic → a cache absorbs the doubling for a fraction of the cost. Long-tail → maybe 2 replicas and a cache. And flag the real wall: at 40% writes and growing, the leader is the next bottleneck, and no replica count fixes writes — that's sharding (Chapter 1.9).
Your photo app has survived everything the last few chapters threw at it. A cache absorbs most reads (chapter 1.6). Replicas mean a dying machine no longer takes you down, and read traffic spreads across the copies (chapter 1.8). Then one morning you look at a graph nobody set an alert on: total database size. 6 terabytes. Growing 400 GB a month.
Here's the part that surprises people: replication cannot help you. Every replica holds a complete copy of the data — five replicas is five machines running out of disk together, in perfect sync. And every write still lands on the one primary; replicas absorb reads, never writes. So you rent a bigger machine. It works. Eight months later you rent an even bigger one, at triple the price. The cloud sells monsters — machines with 24 TB of RAM exist — but the price climbs faster than the spec sheet, and one day you're on the biggest box money rents while your growth curve keeps going. There is no bigger box.
The only move left is the one this chapter is about: stop holding all the data in one place. Cut it into pieces, put the pieces on different machines. The cutting is called sharding, and the cut itself takes one line of code. Everything around the cut is the hard part: which row goes where? Who remembers where each row went? What about queries that need rows from six machines? What happens when you add machine number seven? Each question has a famous wrong answer that has burned real companies. Let's walk through them in order.
There are two ways to slice a database, and they're not equals.
Vertical partitioning cuts by table (or column): users to one machine, photos to another, likes to a third. Companies do this first because it's easy — no routing logic. But look at what it buys: your biggest table still lives on one machine. In your photo app, photos is the monster — 6 TB on its own, growing 400 GB a month. Vertical partitioning moved the monster into its own room; it didn't make it smaller. It buys months, not years — and it quietly kills SQL joins between the tables you separated.
Horizontal partitioning — sharding — cuts by row. Every machine gets the same schema but a different slice of the rows. Add machines, each slice shrinks. This is the cut that scales without a ceiling, and it's what the rest of this chapter is about.
To slice rows across machines you need a rule: this row goes to that shard. The rule is based on one column (or a few) called the shard key — and picking it decides, in one stroke, which queries will be fast and which will be miserable, forever. Any query that includes the shard key routes to exactly one shard: cheap, fast. Any query that doesn't can't be routed — the row could be anywhere, so you must ask every shard. Watch what different choices do to your photo app:
user_id: all of one user's photos live together. "Load user 42's profile" hits one shard. But "fetch photo 987" (you only have the photo ID) could be anywhere.photo_id: "fetch photo 987" hits one shard. But a user's photos are now sprinkled across every shard, so their profile page queries all of them.A good shard key has three properties. High cardinality — many distinct values, so there's something to spread (a country column with 20 values can never spread across 100 machines). Even spread of load — not just row counts; if one value gets most of the traffic, its shard burns while the rest idle. And it must appear in your dominant queries. For most consumer apps that's the entity users act on — user_id, conversation_id, tenant_id — because most requests are "show me my stuff."
The shard key is close to irreversible: changing it means re-deciding where every row lives and moving terabytes while serving live traffic — a months-long migration (chapter 2.11). So when an interviewer asks "how would you shard this?", they're really asking "do you know your dominant access pattern?" Answer that out loud — "reads are 95% 'load this user's data', so I shard by user_id" — and the choice is defensible. Name the key without the access pattern and it's a guess wearing a suit.
In 2011 Instagram was a tiny engineering team with over 10 million users, sharding PostgreSQL. They needed photo IDs that were 64-bit (compact for indexes), sortable by time (feeds show newest first), and generated without new infrastructure — they rejected Twitter's Snowflake ID service because it required running ZooKeeper, more moving parts than their team wanted. Their scheme packed everything into 64 bits: 41 bits of milliseconds since their epoch (January 1, 2011), 13 bits of logical shard ID, 10 bits of a per-shard sequence wrapping at 1,024. IDs are minted inside Postgres itself by a PL/PGSQL function used as the column default — no ID service, and each shard can issue 1,024 IDs per millisecond. The elegant part: every ID carries its own routing — a bit-shift on any photo ID tells you which shard holds it. And their thousands of logical shards (Postgres schemas) mapped onto far fewer physical machines, so growing meant moving whole schemas, never re-cutting rows. Two lessons that generalize: bake the shard into the ID, and create far more logical shards than machines.
You've picked the key. Now: how does a key value map to a shard? Two families.
Range sharding keeps keys sorted and cuts the line into contiguous ranges: keys A–F on shard 1, G–N on shard 2. HBase and Google's Bigtable work this way; MongoDB offers it. The superpower is range queries: "all events between March 1 and March 7" lands on one shard and reads sequentially. The curse is hot ranges. If your key correlates with time — auto-increment IDs, timestamps — every new write targets the last range. One shard sprints while the others sit there like museums of old data. You bought ten machines and got the write throughput of one.
Hash sharding runs the key through a hash function — which turns any input into an effectively random but repeatable number — and uses that to pick the shard. Neighboring keys scatter to opposite ends of the cluster, which is the point: load spreads evenly no matter what the keys look like. The price: ranges are gone. "Photos from last week" means asking everyone, because last week's photos are deliberately everywhere.
| Range sharding | Hash sharding | |
|---|---|---|
| Range scans ("March 1–7") | One shard, sequential read | Every shard — the ordering was destroyed on purpose |
| Load spread | Only as even as your key distribution — time-based keys create one hot shard | Statistically even across shards (per-key hotspots still exist) |
| Adding shards | Split a range in place | Depends on the scheme — this is the next section's drama |
| Used by | Bigtable, HBase, MongoDB (optional) | Cassandra, DynamoDB, Redis Cluster, most hand-rolled app sharding |
The move seniors reach for is the compound key: hash on the entity, range on time within it. Cassandra formalizes this — a partition key (hashed, picks the machine) plus clustering columns (sorted, order rows on that machine). Messages keyed by (conversation_id, sent_at) spread conversations evenly across the cluster, yet one conversation's history is a single sorted run. Both superpowers, one design.
So: hash the key, take it modulo the number of servers — shard = hash(user_id) % 4 — and route. Clean, stateless, one line. Six months later you add a fifth server, and this one line detonates.
Change % 4 to % 5 and recompute where every key belongs. How many keys keep their old home? A key stays only if its hash gives the same answer under both moduli — for 4 to 5 servers, that's about 1 in 5 keys. Eighty percent of your data just changed address. Not because 80% needed to move — the new server only needs a fifth of the data — but because the formula bakes N into every single placement, so changing N invalidates almost everything.
What this does in production depends on what's behind the formula. If it's a cache cluster, 80% of lookups suddenly miss — the data is on the "wrong" node — and the misses stampede your database at once. You've DDoS'd yourself by scaling up. If it's a database, it's worse: cache misses refill themselves, but rows don't teleport. Eighty percent of your rows must physically move over the network before the formula gives correct answers again — and during the move, requests routed by the new formula find nothing at the destination. At terabyte scale, that's not an evening's maintenance window. It's a trap teams have lived in: growth demands a shard, adding a shard demands moving nearly everything, moving nearly everything is unthinkable, repeat.
Mod-N is a school assigning students to classrooms by roll number: room = roll % 3. The day a fourth classroom opens, the school recomputes, and nearly every student in the building switches rooms — even though the new room only needed a quarter of them. The rule looked elegant, but it encoded "3" into every assignment. What you want is a rule where opening a room moves only the students that room will hold.
So the question becomes precise: is there a placement rule where adding a server moves only the data that server takes over — and nothing else? There is, and it's one of the prettiest ideas in distributed systems.
The trick: stop hashing keys onto servers and start hashing keys and servers into the same space. Take your hash function's output range — say 0 to 2³²−1 — and bend the number line into a circle, so the maximum wraps to 0. Hash each server's name onto the circle. Hash each key onto the same circle. The rule: a key belongs to the first server you meet walking clockwise from the key's position.
Now look at what adding a server does. Server D hashes onto the ring — say between A and B. D takes over exactly one thing: the arc of keys between A and D, which used to walk clockwise into B. Every other key still meets the same server it always did. With K keys and N servers, adding one moves about K/N keys — the theoretical minimum, because that's precisely the data the new server must hold to carry its share. Removal is just as tidy: the dead server's arc slides to its clockwise neighbor, and nobody else notices.
Picture a circular road of houses, like numbers on a clock face, with three pizza riders parked at 12, 4, and 8 o'clock. Each house is served by the next rider clockwise. A new rider parks at 2 — and the only houses that change riders are those between 12 and 2, which belonged to the rider at 4. Every other house keeps its rider. In the mod-N school, opening one classroom renumbered the whole city. The analogy's honest limit: real hash functions park the riders at random spots, not neatly spaced ones — and that randomness is the next problem.
Because hashing is random, the arcs are lumpy — with three servers, one can easily land owning half the circle. Worse, think about failure: a dead server's entire arc dumps onto one neighbor, the clockwise successor, which was already carrying a full load. Double load is exactly how one failure becomes two. And there's no way to say "this machine is twice as beefy, give it more."
All three problems share one fix: virtual nodes (vnodes). Place each server on the ring not once but 100 or 1,000 times, at points like hash("serverA-1"), hash("serverA-2"), and so on. Each server now owns many small arcs scattered around the ring. The lumps average out — the same law of large numbers that makes a thousand coin flips come out near 50/50. When a server dies, its hundreds of little arcs are inherited by hundreds of different successors, so its load sprinkles across the cluster instead of crushing one neighbor. A machine with twice the capacity? Twice the vnodes. Cassandra shipped with 256 vnodes per node for years; Amazon's Dynamo paper — the design that popularized all this — used the same idea to keep partitions balanced.
Consistent hashing is everywhere data must spread across a changing set of machines with no central brain: Cassandra and ScyllaDB, DynamoDB and Riak (descendants of the Dynamo paper), memcached client libraries (the Ketama algorithm, built in 2007 precisely so adding a cache node wouldn't flush the fleet), plus load balancers and CDNs. But notice what it optimizes for: automatic, decentralized placement. Some of the most successful sharded systems ever built looked at that property and said — no thanks. I want to know exactly where my data is, and I want moving it to be a decision, not an emergent behavior.
Directory-based sharding abandons formulas entirely. You keep a map — a plain lookup table in a small, strongly consistent store: user 42 → shard 7, tenant "acme" → shard 3, slots 0–511 → machine B. Every request consults the map (in practice a cached copy — the map is tiny and changes rarely) and routes accordingly.
You lose elegance and gain a moving part: the directory must be highly available and never harmfully stale, because a wrong map means reading the wrong shard. What you buy is control. A tenant grows 100x? Edit one row of the map and move just them to a dedicated shard. Draining a machine for replacement? Migrate its entries gradually, at 3pm on a Tuesday, watching dashboards. The mapping is something a human can read, reason about, and fix at 3am — exactly why several famous companies chose it over cleverer schemes.
Mature systems converge on a hybrid: fixed logical partitions plus a movable map. Hash the key into one of many logical buckets — the bucket count never changes, so the mod-N trap never springs — and keep a directory from buckets to machines. Redis Cluster hashes every key into exactly 16,384 slots and maps slots to nodes; resharding means reassigning slots, never re-cutting keys. Instagram's Postgres schemas, MongoDB's chunks with its config-server map, Vitess (the MySQL sharding layer born at YouTube), and Citus all follow the same shape. Rule of thumb: create 10–100x more logical shards than machines on day one. Logical shards are nearly free; re-splitting live data never is.
Pinterest hit hypergrowth in 2011–2012, and their engineers later wrote up the sharding design that carried them through it — a monument to boring technology. They deliberately rejected the era's auto-rebalancing, cluster-managing databases, reasoning that when an opaque balancing algorithm misplaces your data during a failure, you can't fix what you can't understand. Instead: plain MySQL, cut into thousands of small "virtual shards" — each literally its own MySQL database — with the map from shard ranges to physical machines living in configuration, changed only when a human deliberately edits it. Scaling up meant replicating a busy machine, pointing half its shard entries at the replica, and dropping the duplicate halves — whole databases move, individual rows never do. Their 64-bit IDs carry the routing, Instagram-style: 16 bits of shard ID, 10 bits of object type, 36 bits of local ID. Each new user is pinned to a shard where their boards and pins live together, so everyday queries are single-shard local joins. No consistent hashing, no auto-balancer, no magic — and it scaled through one of the fastest growth curves in consumer internet history, run by a small team who could always answer "where is this user's data?" with a config lookup.
So you've sharded — by hash, ring, or directory. The data fits, writes spread, growth has a plan. Now the bill arrives, because every query that doesn't fit the sharding scheme just got expensive.
A query that includes the shard key goes to one shard. Everything else goes to all of them — ask every shard, merge the answers. This is scatter-gather, and its cost is sneakier than "N queries instead of one." The sneaky part is tail latency: your answer isn't ready until the slowest shard replies. Suppose each shard answers slowly (a garbage-collection pause, a disk hiccup) just 1% of the time — a respectable p99. Fan a query out to 100 shards and the odds that at least one is having its bad moment is 1 − 0.99¹⁰⁰ ≈ 63%. Most of your queries now experience somebody's worst 1%. Fan-out turns your shards' p99 into your users' median.
The senior mitigations, in order of preference: don't scatter — design so hot queries carry the shard key; build a lookup path — if users log in by email but you shard by user_id, keep a mapping table (email → user_id), itself sharded by email, so login is two single-shard hops instead of a broadcast (DynamoDB's global secondary indexes are this idea productized); and for the scatter you can't avoid, hedge — send a backup request to a replica when a shard is slow. Note what fell off the list: making shards faster. At 100-way fan-out, even great shards produce bad tails. You fix the fan-out, not the shards.
Two more line items. Joins across shards effectively don't exist — you denormalize (store copies where they're read) or join in application code. Uniqueness gets weird: with users on 50 shards, "is this username taken?" has no single place to ask — unless you make one, a reservation table sharded by username itself. And the biggest item deserves its own paragraph.
Cross-shard transactions. Move ₹500 from account A on shard 1 to account B on shard 4. Debit A, commit. Before the credit lands, the app server crashes. The money is gone — not stolen, just nowhere: debited, never credited. Within one machine, the database's transaction machinery makes "both or neither" a guarantee. Across two machines, no single database can promise it, because no single database saw both writes. Solving this properly — two-phase commit, sagas, outboxes — is hard enough to be chapter 2.1. The sharding-level defense is to need it rarely: choose a shard key that keeps transactions single-shard. It's why Pinterest co-locates a user's data, and why financial systems obsess over what lives with what.
One more failure mode — the one perfect hashing can't touch. Hashing spreads keys evenly. It does nothing about the load per key, and one key can be a hurricane: a celebrity's profile, a viral post, a Discord channel with 500,000 members, a B2B tenant 1,000x your median customer. Hash it beautifully; it still lands, whole, on one shard, and that shard melts while nineteen others nap.
This is the hot partition (or hot key, or celebrity) problem, and it gets a full chapter (2.6). The preview: cache the hot thing in front of the shard, split the entity itself (append a suffix — bieber-1 through bieber-16 — spreading one giant key over many shards, at the cost of gathering on read), or isolate it — with a directory, moving the whale to its own shard is a one-line map edit, a big reason directories keep beating elegant math in real companies.
In 2017 Discord moved its messages to Cassandra, partitioned by (channel_id, bucket) — the bucket a ten-day time window, so one channel's history is chopped into slices of time rather than one endless partition. Sensible design. By 2022 the cluster held trillions of messages across 177 nodes and had become, in their own telling, an on-call nightmare — largely from hot partitions. When something big happened in a huge channel, traffic concentrated on the one partition holding that channel's current bucket; the node holding it fell behind, and latency rippled outward to unrelated channels on the same node, with Cassandra's JVM garbage-collection pauses amplifying every incident. The 2023 fix came in two parts. They migrated to ScyllaDB — a C++ rewrite of Cassandra, no garbage collector, shard-per-core — same data model, calmer engine. But the deeper fix was architectural: Rust-based "data services" between API and database doing request coalescing — a thousand users ask for the same hot message page, one query reaches the database, everyone shares the answer. A Rust migrator moved trillions of messages in about nine days; the fleet shrank from 177 nodes to 72, and message-read p99 fell from 40–125 ms to around 15 ms. The interview lesson: a hot key is not a sharding problem, so more shards can't fix it — you need coalescing, caching, or isolation in front of the shard.
Zoom out and the industry's choices form a pattern worth memorizing:
| System | Scheme | Worth knowing |
|---|---|---|
| Cassandra / ScyllaDB | Consistent hashing + vnodes | Compound keys: hashed partition key picks the node, clustering columns sort within it |
| DynamoDB | Hash partitioning, automatic splits | Descends from the Dynamo paper; splits hot partitions adaptively |
| MongoDB | Range or hash chunks + directory | Config servers hold the chunk map; a balancer migrates chunks between shards |
| Redis Cluster | 16,384 fixed hash slots + slot map | The fixed-logical-buckets pattern in its purest form |
| Vitess / Citus | Logical shards + directory | Sharding layers over MySQL/Postgres; Vitess grew up inside YouTube |
| Kafka | hash(key) % partitions | Literal mod-N! Partition count is meant to be fixed — adding partitions remaps keys and breaks per-key ordering, so teams over-provision partitions up front |
Notice two things. First, nobody re-cuts rows: mature systems move whole logical units — vnodes, slots, chunks, schemas. Second, Kafka proves mod-N isn't evil, it's conditional: with a fixed N it's the simplest correct scheme there is. The disaster is only ever mod-changing-N.
Sharding is a favorite senior probe because the first answer is easy and every follow-up has teeth. A passing senior answer names the shard key and the dominant query that justifies it, then volunteers the costs — cross-shard queries, hot keys, resharding — before being asked. Expect these:
Explain consistent hashing to a junior engineer in five or six sentences — including why plain hash % N fails and what virtual nodes are for.
When you spread data across N servers with hash(key) % N, the number N is baked into every placement — so the moment you add one server, almost every key's answer changes and nearly all your data has to move. Consistent hashing fixes this by hashing both servers and keys onto the same circle of numbers, and giving each key to the first server clockwise from it. Now adding a server only claims the small arc between it and its neighbor — roughly 1/N of the keys move, which is the minimum possible, and nothing else is touched. One catch: random placement makes the arcs uneven, and a dead server dumps its whole arc on one neighbor. So each real server is placed on the ring hundreds of times as "virtual nodes" — the many small arcs average out to a fair share, and when a server dies its load sprinkles across everyone instead of crushing one machine. That's why Cassandra, DynamoDB, and most distributed caches use exactly this.
You're sharding the messages table for a chat app (think WhatsApp: 1:1 and group chats). Candidate shard keys: message_id, sender_id, conversation_id. Pick one, justify it with the dominant query, and name the failure mode your choice keeps.
conversation_id. The dominant query is "load recent messages of this conversation" — every chat open, every scroll — and this key makes it one shard, one sequential read (add a time-ordered clustering column). sender_id scatters every conversation across shards; rendering one chat becomes a multi-machine merge-sort. message_id is worst: perfectly uniform, and every real query becomes scatter-gather. The kept failure mode: a mega-group is a hot partition on one shard — exactly Discord's problem — so you'd add time-bucketing to the key and request coalescing in front.
Your memcached tier has 4 nodes addressed by hash(key) % 4, running at a 95% hit rate. You add a fifth node. Roughly what fraction of keys change nodes, what happens to the hit rate in the next minute, and who feels it? Then: same event with consistent hashing — what changes?
With mod-N, a key keeps its node only when h % 4 == h % 5 — about 1 in 5 keys — so ~80% change address and miss. Hit rate crashes from 95% toward ~20% instantly. Who feels it: the database, which was serving 5% of reads and suddenly serves ~80% — a 16x read spike you inflicted on yourself, possibly an outage. With consistent hashing, the new node takes ~1/5 of the keyspace; hit rate dips to ~76% and recovers as the node warms. This exact scenario — cache fleets flushing themselves on resize — is why Ketama was built for memcached in 2007.
A teammate proposes sharding a URL shortener's links table by created_at range: January's links on shard 1, February's on shard 2, and so on. "Range sharding, clean splits, easy to add shards." What breaks?
Both hot spots at once. Writes: every new link has this month's timestamp, so 100% of inserts hammer the newest shard — N machines, one machine's write throughput. Reads too, because click traffic decays with link age: the newest shard takes nearly everything while old shards store cold history at full hardware cost. The correct move is hashing the short code — point lookup by exact code is the only query that matters, and range scans over creation time are worthless to users, so range sharding's price buys a superpower nobody uses. Time-correlated keys plus range sharding is the conveyor-belt trap: one worker at the end of the belt doing everything.
Capacity planning: your photos table is 6 TB, growing 400 GB/month, and you're comfortable running shards up to ~2 TB. How many shards do you create, and on how many machines? (Hint: the number of shards and the number of machines should not be the same number.)
The junior answer is "3 shards, maybe 4 for growth" — which schedules the next painful reshard about a year out. The senior answer separates logical from physical: create ~64 logical shards on day one (each ~95 GB — small is fine, logical shards are nearly free) mapped onto 4 machines, 16 shards each at ~1.5 TB per machine. When a machine nears its limit, add a machine and move whole logical shards to it — bulk copy plus map update, no row-by-row re-cutting, ever. 64 logical shards buys headroom to 64 machines (~128 TB) before a true re-split. This is the Instagram/Pinterest/Redis-Cluster pattern: fix the bucket count high, move buckets.
hash % N moves ~all data when N changes; consistent hashing moves only ~K/N keys — the minimum — and virtual nodes fix its lumpy arcs and spread a dead node's load.It's the biggest sale night your e-commerce app has ever seen. You run two data centers — Mumbai and Singapore — each holding a full copy of the data, so that losing one building doesn't mean losing the company. At 11:52 pm, an ISP botches a router update, and the two data centers can no longer reach each other. Both are alive. Both are serving users. They just can't talk.
At 11:53, a user named Priya taps "Add to cart" on a pair of headphones. Her request lands in Mumbai. Mumbai cannot check with Singapore, and cannot replicate the write to it. You have exactly two options, and your system will pick one in the next 200 milliseconds whether you designed for it or not. Option one: refuse — show Priya an error, because accepting a write Singapore can't see means the two copies of her cart now disagree. Option two: accept — take the write in Mumbai, let the copies disagree, and clean up the mess after the network heals.
There is no third option. "Wait for the network to come back" is option one with extra steps — the request times out, Priya sees a spinner, and buys her headphones somewhere else. Every system that keeps copies of data in more than one place faces this exact fork. Everything in this chapter — CAP, consistency levels, quorums, conflict resolution — is machinery for choosing a side of that fork on purpose, instead of by accident at midnight.
You've probably seen the CAP triangle: Consistency, Availability, Partition tolerance — "pick any two!" — as if the three were flavors on a menu. That framing has confused a generation of engineers, because it suggests "partition tolerance" is optional, something you can trade away for the other two. It isn't. Let's define the words plainly first:
Here's the fix for the misunderstanding: partitions are not a choice you make; they're a thing that happens to you. Switches die. Fiber gets cut by a backhoe. A garbage-collection pause freezes a node for ten seconds, and to everyone else that's indistinguishable from a partition. A routine config push blackholes traffic between racks. The first of the famous fallacies of distributed computing is "the network is reliable" — it's listed first because it's the one everyone falls for. If your data lives on more than one machine, a partition will occur.
So the honest reading of CAP is: when a partition happens, you must choose between consistency and availability. That's it. Two options, chosen only during the partition:
What about "CA"? A system that is consistent and available but not partition-tolerant is a system where the network never splits — which means a single machine, or a fantasy. The moment you replicate across a network, "CA" stops being on the menu.
One more correction that matters at the senior level: CP and AP are not zodiac signs for databases. The choice is made per operation, during a partition. The same system can refuse checkout writes (CP behavior) while happily serving slightly stale product pages (AP behavior). Hold that thought — it's the punchline of the whole chapter.
Two branches of a coffee shop share gift-card balances by phoning each other before every sale. One day the phone line dies. The CP branch says "sorry, card machine's down" — it refuses sales it can't verify, stays perfectly correct, and loses angry customers. The AP branch keeps swiping cards and jots sales in a paper ledger to reconcile tonight — customers are happy, but a clever one can now spend the same ₹500 balance at both branches. Refuse and stay right, or serve and clean up later: that's the entire theorem. Where the analogy breaks: real systems reconcile automatically and constantly, not once at closing time — and doing that reconciliation well is its own section below.
During routine maintenance — replacing some failing optical equipment — GitHub lost connectivity between its US East Coast data center and its network hub for just 43 seconds. That was enough. Orchestrator, their automated MySQL failover tool, saw the East Coast primaries vanish and did its job: it promoted replicas on the West Coast and pointed all write traffic there. But replication between coasts was asynchronous, and the East Coast primaries had accepted a few seconds of writes that had never made it West. When the 43-second partition healed, GitHub had two divergent primaries — a split brain: writes on the East that the West had never seen, and a growing pile of new writes on the West (applications wrote there for roughly 40 minutes before engineers fully understood the topology). There was no automatic way to merge them. GitHub chose consistency for the recovery: they degraded the site for more than 24 hours — pausing webhooks and Pages builds, restoring from backups, replaying replication, and reconciling the orphaned writes by hand. The sharpest lesson: their failover automation had silently made an AP choice across a boundary where the data could not tolerate divergence. Nobody had decided that; the tooling had. CAP doesn't care whether you made the choice deliberately — it only guarantees you made one.
Here's the thing that makes CAP feel academic: real partitions are rare. A well-run network might see minutes of true partition per year. If the consistency trade-off only mattered during those minutes, you could almost ignore it. But there's a second trade-off hiding in plain sight, and you pay it on every single request. PACELC names it: if there's a Partition, choose Availability or Consistency — Else, choose Latency or Consistency.
Why is there a choice even when the network is healthy? Physics. If your replicas are in Virginia and Frankfurt, a network round trip between them costs roughly 80–100 milliseconds — the speed of light doesn't negotiate. To be strongly consistent, a write must wait for far-away replicas to confirm before telling the user "saved." That's your write latency: local disk write plus a transatlantic round trip. Or you can acknowledge the write locally in 5 milliseconds and replicate in the background — fast, but now a reader in Frankfurt can see data that's a second stale. Strong consistency isn't just fragile during partitions; it's slow every day. That everyday tax, not the rare partition, is usually why teams relax consistency.
The vocabulary is worth thirty seconds in an interview: DynamoDB and Cassandra are PA/EL systems — available during partitions, fast otherwise. Google Spanner is PC/EC — it refuses rather than diverge, and pays coordination latency on every transaction because Google decided its databases should never lie, and bought atomic clocks to keep the tax small. Neither is "better." They charge different bills to different customers.
"Strong versus eventual" sounds like a binary. It's actually a staircase, and each step down buys latency and availability by permitting a specific, describable user-visible bug. The fastest way to internalize the levels is to learn each one as the bug report that arrives when it's missing.
Linearizable (strong): the system behaves as if there's one copy of the data. The instant a write completes, every reader everywhere sees it. No bug reports — just the latency and availability bill from the previous section. You buy this for money movement, inventory decrements, username uniqueness.
Eventual consistency: the default of the internet. The only promise: if writes stop, all replicas converge to the same value, eventually — usually within milliseconds to seconds. Meanwhile, any read may be stale. The bug report: "the like count says 41, I refresh, it says 40, I refresh, 41 again." For a like count, nobody files that bug. That's the point — for most data, eventual is fine and nobody notices.
Read-your-writes: a targeted upgrade — you always see your own writes; other people's may lag. The bug when it's missing is the most famous one in replication: a user, Priya, posts a comment, the write goes to the primary database, her next page load reads from a replica that's two seconds behind, and her comment is gone. She posts it again. Now there are two. Users don't file this as "replication lag" — they file it as "your app ate my comment," and they're right. Common fixes: pin a user's reads to the primary for a few seconds after they write, or have the client remember the version it wrote and demand at-least-that-fresh reads.
Monotonic reads: time never runs backwards for one reader. The bug when it's missing: Priya's first refresh hits a fresh replica and shows a friend's comment; her second refresh hits a more stale replica and the comment disappears again. Seeing data appear and then un-appear feels haunted. Fix: keep each session sticky to one replica, so a reader's view only ever moves forward.
Causal consistency: effects never appear before their causes — a reply never shows up before the message it answers. The bug when it's missing is genuinely damaging: in a group chat, Bob's "No, absolutely not!" gets replicated faster than Alice's question "Should I delete the production database?", and a third user sees the answer before the question. Reads make sense only in order. Causal consistency tracks the "happened because of" relationship and delivers writes in an order that respects it — noticeably cheaper than linearizability, because unrelated writes can still flow freely.
| Level | The promise | Bug report when it's missing | Relative cost |
|---|---|---|---|
| Linearizable | Everyone sees every write instantly, like one copy | — | Coordination on every operation; cross-region latency; reduced availability |
| Causal | Effects never precede causes | "I saw the reply before the question" | Metadata to track what-caused-what |
| Read-your-writes | You always see your own writes | "Your app ate my comment (so I posted it twice)" | Session stickiness or version tokens |
| Monotonic reads | One reader's view never goes backwards | "The comment appeared, then vanished" | Sticky sessions |
| Eventual | Replicas converge if writes stop | "The count flickers between 40 and 41" | Nearly free — the internet's default |
Notice that read-your-writes and monotonic reads are cheap, targeted patches: they don't make the system globally consistent, they just kill the two most embarrassing bug reports. A senior answer names the level per feature instead of reaching for "strong everywhere" — because now you know what each level costs and exactly which bug each level prevents.
So far "replication" has been vague. Here's the concrete machine that lets you tune where you sit on the staircase. Say every piece of data is stored on N replicas (N=3 is the industry default). For a write, you wait for W replicas to confirm before telling the user "saved." For a read, you ask R replicas and take the answer with the newest version number. N, R, W are knobs you set.
The magic inequality: R + W > N. If the write touched W nodes and the read asks R nodes, and R+W exceeds N, then the two groups are too big to avoid each other — at least one node you're reading from must be one of the nodes that confirmed the latest write. Version numbers on each copy tell you which answer is the fresh one. It's pigeonhole arithmetic, nothing deeper.
You keep your class notes in three notebooks held by three friends: Asha, Bilal, and Chen (N=3). You're too lazy to update all three, so your rule is: every update goes into any two notebooks (W=2), and before an exam you always check any two notebooks (R=2) and trust whichever page has the later date. Pick any two to write and any two to read — the groups must share at least one friend, because there are only three friends. You can never check two notebooks and have both be out of date. That's R + W > N: 2+2>3. The analogy breaks in one honest place: real replicas can fail mid-update, and the date-comparison has sharp edge cases — quorum overlap gets you "reads see the latest completed write" in the common case, not a full mathematical guarantee of linearizability. Databases layer extra machinery (like read repair) on top.
The knobs let you buy exactly what you need:
| N / W / R | What you bought | What you paid |
|---|---|---|
| 3 / 2 / 2 | Overlap guarantee; tolerates one node down for both reads and writes | Every operation waits on two nodes — the balanced default |
| 3 / 3 / 1 | Blazing one-node reads, still overlapping — great for read-heavy data | One dead node and all writes fail |
| 3 / 1 / 3 | Instant writes | Slow reads, and a write confirmed by one node can die with that node — durability risk |
| 3 / 1 / 1 | Maximum speed and availability | R+W = 2 ≤ 3: no overlap, no freshness promise. This is eventual consistency, chosen honestly |
One more twist. A strict quorum refuses a write when the data's designated home nodes are unreachable — CP behavior. A sloppy quorum says: if a shopping cart normally lives on nodes A, B, C and node C is unreachable, write to substitute node D instead, and attach a note — "this belongs to C; deliver it when C comes back." That note-and-return mechanism is called hinted handoff. Availability goes way up: writes succeed as long as any W nodes are alive, anywhere. The price: while the hint sits on D, your overlap math is broken — a read of A, B, C can miss the newest write because it's living at D's house. Sloppy quorums are an explicit AP move: never refuse the write, tolerate a window of missable data.
Choose availability — or sloppy quorums, or async replication — and eventually the inevitable happens: two replicas accept different writes to the same key, and then meet. Both have version "8" of the same cart, and they disagree. Someone has to decide what the truth is. There are three honest strategies, and one of them is a trap painted like a strategy.
Last-writer-wins (LWW): stamp every write with a timestamp; when copies conflict, keep the latest stamp, discard the rest. It's simple, automatic, and let's say the quiet part loudly: LWW silently deletes data your system already acknowledged. The "losing" write was accepted, confirmed to a user, and is now gone — no error, no log line, no trace. Worse, the winner is chosen by server clocks, and server clocks disagree: even machines that sync over NTP (the internet's clock-synchronization protocol) commonly drift milliseconds apart, and badly configured ones drift seconds. A node with a fast clock wins conflicts with writes from the "future." Cassandra resolves conflicts this way by default, and Jepsen — an independent test project that tortures databases with induced partitions and publishes the wreckage — famously demonstrated it dropping acknowledged writes under contention. LWW is a fine choice when losing a conflicting write costs nothing — a device's "last seen location," a cache entry. It is a disaster for carts, counters, and documents, and choosing it there is choosing silent data loss.
Vector clocks: the diagnostic upgrade. Concept only, no math: each replica keeps a tally of how many writes it has seen from every replica, and each version of the data carries that tally. Comparing two versions' tallies answers one question: did one version already know about the other when it was written? If yes, the newer one safely supersedes it — no conflict, keep the descendant. If neither knew about the other, they're concurrent — a genuine conflict, and the system now knows it instead of guessing. That's the crucial difference from LWW: vector clocks detect conflicts honestly; they don't resolve them. Something still has to merge.
Application-level merge: the only resolver that knows what the data means. The database sees two mystery blobs; your application knows they're shopping carts, and that the sane merge of two carts is the union of their items. Two concurrent counter increments? Add them. Two edits to a document? That's the hard end — operational transforms and CRDTs (special data structures built so concurrent edits always merge to the same result) live there. Merge logic is real work, which is why it's reserved for data that deserves it.
During the 2004 holiday shopping season, Amazon suffered a string of outages traced to pushing their Oracle-based infrastructure past its limits — clustering and replication on the leading commercial database of the day simply couldn't sustain the availability and scale the retail business demanded. The post-mortem question was blunt: what does the shopping cart actually require? Answer: it must always accept adds. Every refused "add to cart" during a partition is a customer walking away with their wallet — availability is literally revenue. So Amazon built Dynamo, a key-value store that inverted the traditional priorities: always-writable via sloppy quorums and hinted handoff, vector clocks to detect concurrent versions, and conflicts surfaced at read time for the application to merge — the cart merges by union of items. The paper is honest about the cost: after a merge, a deleted item can occasionally resurrect in your cart, and Amazon judged "a deleted item reappears (rarely)" a fair price for "adds never fail." The 2007 paper became one of the most influential systems papers ever written — Cassandra, Riak, and Voldemort are its direct descendants. And the epilogue completes this chapter's argument: DynamoDB, the 2012 managed service, offers both — eventually consistent reads by default at half the cost of strongly consistent ones, strong reads when you ask, and last-writer-wins for its cross-region global tables. Even the company that wrote the availability manifesto ended up selling consistency as a per-request checkbox.
Now assemble the whole chapter into the sentence that separates senior candidates from everyone else. The junior move is picking a side: "I'd use an AP database because availability matters." The senior move recognizes that one system contains many kinds of data, and each gets its own contract. For an e-commerce site: the cart is AP with union-merge (a refused add costs a sale; a merge conflict costs nothing). Checkout and payment are CP behind a quorum (a duplicate charge or oversold last-item costs money and trust — better to show "please retry" for thirty seconds a year). Product pages are eventual with heavy caching (a price that's ninety seconds stale is fine; validate the real price at checkout). Order history gets read-your-writes (users must see the order they just placed). Same product, four consistency contracts — each one an explicit decision with a reason attached.
Consistency questions at the senior level are commitment traps: the interviewer offers you a chance to recite dogma, and grades you on refusing it. A passing answer sounds like: "During a partition, cart writes keep flowing — union-merge on read, worst case a deleted item resurfaces. Inventory decrement is quorum-backed and will refuse writes on the minority side; I'll show 'try again' there because overselling is worse than an error. And day-to-day the real trade is latency: I'll replicate to the second region async and accept a one-second staleness window, because 100ms synchronous writes would show up on every purchase." Probes to expect:
A junior teammate says "we should pick two of CAP's three letters — I vote consistency and availability." Correct them in five sentences, kindly.
Partitions aren't one of the choices — they're the weather; the network will split whether we vote or not. CAP really says that when a partition happens, each operation must pick one of two behaviors: refuse requests it can't verify (consistency) or answer anyway and reconcile the copies later (availability). "Consistent and available with no partitions" just describes a single machine, which is not what we're building. Even on healthy days there's a milder version of the same trade: waiting for far-away replicas makes every write consistent but slow, so latency versus consistency is the choice we actually pay daily. So instead of voting once for the whole system, we should decide per feature — refuse-and-retry for payments, accept-and-merge for the cart.
Four bug reports land in your queue. Name the missing consistency guarantee for each: (a) "I posted a review, refreshed, and it was gone — so I posted it twice." (b) "My friend's comment showed up, then disappeared on the next refresh, then came back." (c) "In the group chat I saw 'No, cancel it!' before I saw the question it was answering." (d) "The product's like count wobbles between 128 and 131."
(a) Read-your-writes — her write went to the primary and her read hit a lagging replica; fix by pinning her reads to the primary briefly after a write. (b) Monotonic reads — successive reads hit replicas with different lag, so her view of time went backwards; fix with sticky sessions to one replica. (c) Causal consistency — the reply replicated faster than its cause; fix by tracking and respecting the happened-because-of order. (d) Just eventual-consistency staleness — different replicas at slightly different versions. This one is arguably not a bug worth fixing: the cost of making a like count linearizable exceeds anyone's benefit. Recognizing which report to close as "working as intended" is part of the skill.
You run a metrics dashboard with N=5 replicas per shard: about 50,000 reads/sec and 500 writes/sec. Choose W and R. Then answer: what happens to reads and writes when two of the five nodes die?
Read-heavy by 100:1, so make reads cheap: R=1, and keep the overlap by setting W=5? No — W=5 means one dead node kills all writes. The sweet spot is W=4, R=2 (4+2>5) or the balanced W=3, R=3. With W=4/R=2: reads touch just two nodes (cheap, and this system does 100x more reads), writes touch four (fine at only 500/sec). Now kill two nodes: three remain. Reads (R=2) still work. Writes need W=4 confirmations from three living nodes — all writes fail. If that's unacceptable, W=3/R=3 survives two dead nodes for both operations, at the price of reads costing three nodes each. The exercise's real lesson: R and W tune not just latency but which failure you'll experience — say the failure mode out loud when you pick the numbers.
Your team stores collaborative to-do lists in a Dynamo-style store with LWW conflict resolution. App server clocks drift up to 30 seconds apart (someone misconfigured NTP on half the fleet). Two users edit the same list from different servers within the same minute. What exactly goes wrong, what does the user experience, and what are two fixes?
Both edits are acknowledged. When the replicas reconcile, LWW compares timestamps — but the timestamps measure clock configuration, not reality. The server with the fast clock wins even if its write happened first in real time; the other user's acknowledged edit is silently discarded — no error, no log, just a to-do item that un-happens. The user experience: "I added 'buy milk,' the app said saved, and an hour later it's gone." These are the worst bugs to debug because nothing failed — the system did exactly what LWW specifies. Fix one: switch conflict handling to detect-and-merge — vector clocks to identify concurrent versions, union of list items as the merge (this is exactly the Dynamo cart design). Fix two (mitigation, not cure): fix NTP and monitor clock skew as a paged alert — but say plainly that better clocks shrink the window and never close it; two writes in the same millisecond still tie. For shared mutable data, LWW is the bug.
Sketch the consistency contract, component by component, for a food-delivery app: (a) restaurant menu pages, (b) the user's cart, (c) placing the order and charging the card, (d) live driver location on the map, (e) "your order status" screen. For each: level, CP-or-AP during a partition, and the one-line reason.
(a) Menus: eventual, AP, cached hard — a menu 60 seconds stale hurts nobody; validate price at order time. (b) Cart: AP, always-writable with union merge — a refused add loses the sale; a merged cart is at worst mildly surprising. (c) Order + charge: CP, linearizable, quorum-backed — a double charge or an order accepted by a kitchen that never sees it costs real money and trust; "please retry" for a rare partition window is the cheaper harm. (d) Driver location: eventual with LWW — and here LWW is correct: each new ping fully supersedes the old one, so discarding a conflicting stale ping loses nothing. (e) Order status: read-your-writes — the user must see the order they just placed, or they'll place it again (which, combined with (c), is how double orders happen). The senior tell in this answer is (d): knowing when LWW is fine matters as much as knowing when it's a trap.
A user taps Place Order in your shop. Your handler gets to work: charge the card (300 ms), write the order row (20 ms), call the email provider for a confirmation (800 ms), render a PDF invoice (2 seconds), ping the warehouse (300 ms), update analytics (100 ms). Then, finally: "Order confirmed!" About 3.5 seconds of spinner. Your users grumble, but they cope.
Then one Tuesday your email provider has a bad day. Their API slows to 30 seconds, then starts timing out. And suddenly checkout is down. Cards are charging fine. Your database is bored. But every order request dies waiting on a confirmation email — a thing nobody on Earth needs within the next second — and the user sees a red error banner over an order that mostly succeeded. You are losing real money because a courtesy email is slow.
Look at what you actually built: you made the user wait for work they can't see, and you chained your availability to your flakiest dependency. The fix is one of the oldest moves in system design: do what must happen now, promise the rest, and keep that promise a few seconds later. The machine that holds the promise is a message queue, and this chapter is about using one like an adult — including the sharp edges the marketing pages skip.
Take a checkout handler that charges the card, writes the order row, calls an email provider for a confirmation, renders a PDF invoice, pings the warehouse and updates analytics before it replies. Every piece of that work is either synchronous — the user's response depends on its result — or asynchronous — it must happen, but not before you reply. The test: if this step silently took ten minutes, would the response I just sent become a lie?
So the handler becomes: charge, write, drop four small job messages onto a queue, reply. Run one after another those six steps cost the user about 3.5 s of spinner; now the reply comes back in ~350 ms. And when the email provider melts — its API slowing to 30 seconds, then timing out — checkout doesn't notice, because the email jobs just sit in the queue until it recovers.
One caveat before we celebrate. "Async" means eventual, and eventual raises new questions: what if the email job fails? What if it runs twice? How do I even know it ran? The rest of this chapter is those answers — "throw it on a queue" is only the easy half.
A message queue has three roles. A producer writes a small message — "send confirmation for order 8213" — and moves on. A broker (the queue system itself) stores that message durably, on disk, usually replicated, so a crash doesn't eat it. A consumer (a worker) pulls messages and does the actual work at its own pace. Producer and consumer never talk directly, and that gap is the whole point.
The gap buys three kinds of decoupling. Time: the consumer can be down for ten minutes and nothing is lost — messages wait. Rate: producers can burst at 50,000 messages a second while consumers steadily chew through 5,000; the queue holds the difference. Failure: a sick consumer no longer makes the producer's request fail — exactly what keeps the checkout handler above alive when its email provider melts.
A restaurant kitchen runs on a queue: the ticket rail. The waiter (producer) clips an order to the rail and goes straight back to the dining room — they don't stand at the stove until the biryani is done. Cooks (consumers) pull tickets at the pace the stove allows. When the 8pm rush hits, the rail fills instead of the kitchen collapsing; when a cook steps out, tickets wait instead of vanishing. Add a cook, and the rail drains faster — no waiter changes anything. Where the analogy breaks: a rail is small and visible, so overload is obvious. A real queue can silently grow to millions of messages — which is why monitoring queue depth becomes non-negotiable later in this chapter.
Queues come in two shapes, and mixing them up causes real bugs.
Point-to-point (a work queue): many consumers may be attached, but each message is processed by exactly one of them — they compete for work, like cooks pulling from one rail. This is the shape for jobs: send this email, resize this image. Done once, by whoever's free.
Publish/subscribe (pub/sub): the producer publishes an event — "order 8213 was placed" — to a named topic, and every subscriber gets its own copy: email, warehouse, analytics each receive one. This is fan-out: one fact, many independent reactions.
The deep difference is who knows what. In pub/sub, the producer doesn't know or care who's listening. Six months from now, the fraud team can subscribe to order.placed and start scoring orders — without a single line of the order service changing. You post facts on a bulletin board; teams you've never met build on them.
Around 2010, LinkedIn's data plumbing was drowning in exactly the problem pub/sub solves. Activity data — page views, clicks, profile visits — and operational metrics had to flow from many source systems into many destinations: the warehouse, Hadoop, search indexes, monitoring, recommendations. Each connection was its own hand-built, point-to-point pipeline, so the pipeline count grew like sources × destinations, and every new destination meant another fragile custom feed. Jay Kreps, Neha Narkhede, and Jun Rao replaced the tangle with one idea: a distributed, partitioned, replicated commit log. Every event is appended once to a durable, ordered log; any consumer reads at its own pace, tracking its own position — even rewinding to replay history. Publish once, consume anywhere; the N×M mess became N+M. Open-sourced in 2011, and by 2019 LinkedIn's deployment carried around seven trillion messages a day. The name? Kreps had taken a lot of literature classes, and since it was "a system optimized for writing," he named it after a writer: Franz Kafka. Chapter 3.17 has you build a mini version yourself.
Here's the question that separates people who've used queues from people who've read about them: when a consumer crashes halfway through a message, what happens?
Walk through the mechanics. A consumer pulls "send email for order 8213." It sends the email. Then, a millisecond before it tells the broker "done" — the machine dies. The broker now faces an unsolvable puzzle: did the consumer crash before the work, or after the work but before confirming? Both look identical from outside: a message went out, no confirmation came back. The broker must guess, and there are only two guesses:
At-least-once is the default reality of every serious queue system, because for most work a duplicate is annoying but a silent loss is a bug you'll never find. Which forces the follow-up: what about exactly-once, the thing every broker's landing page advertises?
The blunt truth: exactly-once delivery across a network is not achievable — the crashed-before-or-after puzzle has no answer, so duplicates (or drops) are always possible at the delivery layer. What actually works, and what honest "exactly-once" marketing means, is at-least-once delivery plus idempotent processing: handling the same message twice has the same effect as once. Check "did I already send email for order 8213?" before sending; use the message ID as a database key so a second insert bounces off. Kafka's genuine exactly-once feature is real but scoped — transactions within Kafka-to-Kafka pipelines; the moment your consumer touches the outside world (an email API, a card charge), you're back to needing idempotency. Chapter 2.2 builds that properly. For now, the senior reflex: assume every message can arrive twice, and make twice harmless.
Inviting a friend to a party over a bad phone line. At-most-once: say it once, hang up — maybe they heard. At-least-once: repeat until they say "got it" — but if the line drops right after they heard you, you'll call back and they'll hear it twice. Exactly-once over a bad line is impossible to guarantee — silence could mean "didn't hear" or "heard, reply lost." So you make hearing it twice harmless: "Party, Saturday, invite #42" — your friend checks their list, sees #42 already noted, and just says "got it" again. That check is idempotency, and it's the real machinery behind every exactly-once promise.
The confirmation a consumer sends is an acknowledgment (ack): "done, delete this message." At-least-once falls out of one rule — ack only after the work is finished. (Ack-before-work is how you accidentally build at-most-once.)
But if a consumer takes a message and dies without acking, how long should the broker wait before concluding it's dead? That's the visibility timeout: a received message isn't deleted — it becomes invisible to other consumers for a window, say 60 seconds. Ack within the window and it's deleted for good; miss it and the message reappears for someone else to try. A library loan: the book is off the shelf while you have it, and if you never return it, the library declares it lost and orders another copy.
Sizing this number is a real operational decision. Too short — 30 seconds on a 2-minute job — and the broker redelivers messages still being processed: two workers do the same job, duplicates manufactured out of thin air. Too long — 30 minutes — and a genuinely crashed consumer's message waits 30 minutes for a retry. Rule of thumb: comfortably above your slowest normal processing time, with the broker's heartbeat API extending the window for long jobs genuinely in progress.
Kafka, being a log, does acks differently: consumers periodically commit an offset — a bookmark saying "everything up to position 4,181 is done." Crash, and you resume from the last committed bookmark, replaying whatever came after. Different mechanism, same honest guarantee: at-least-once, duplicates possible, idempotency required.
Now a nastier failure. A message arrives with a malformed payload — a producer bug put a null where the email address should be. Your consumer pulls it and crashes. The visibility timeout expires, the message reappears, another consumer pulls it and crashes. Forever. This is a poison message, and left alone it's a tiny denial-of-service attack from inside your own system: it burns workers in a crash loop, and on a strictly-ordered queue it blocks everything behind it (head-of-line blocking) while it retries.
The standard defense is a dead-letter queue (DLQ). The broker counts deliveries per message; after N failed attempts — five is a common default — it stops retrying and shunts the message to a separate holding queue. The main queue flows again, and the poison sits in quarantine with its delivery history, waiting for a human.
Two operational habits make a DLQ actually useful, and both are senior signals. First, alert on it: a DLQ with no alarm is where messages go to be forgotten — page someone when its depth goes above zero. Second, have a redrive plan: after fixing the bug, replay DLQ messages into the main queue so those customers eventually get their emails. Dead-lettered messages are deferred work, not deleted work.
With many consumers in parallel, messages finish out of order — worker A gets message 1, worker B gets message 2, B finishes first. Usually nobody cares: it doesn't matter whose confirmation email sends first. But sometimes order is the whole point. A bank account that processes "withdraw 500" before "deposit 1,000" bounces a payment that should have cleared.
Here's the trade to state crisply at a whiteboard: a total, global order means processing messages one at a time — one consumer, no parallelism. Global ordering and horizontal scaling are directly opposed. The escape is noticing that nobody actually needs global order. The bank needs order per account; a chat app, per conversation. Order within a key, stay parallel across keys.
Kafka makes this concrete with partitions. A topic is split into partitions, each an independent ordered log. The producer hashes a key — account_id, say — to pick the partition, so every message for account 42 lands in the same partition, in order, while different accounts spread out and process in parallel. SQS FIFO queues offer the same idea as "message groups": strict order within a group, parallelism across groups — with the caveat that per-group throughput is limited to hundreds-to-thousands of messages a second, far below unordered queues. Ordering is never free; buy it only where a key genuinely needs it.
Scaling consumers on a work queue is beautifully boring: queue's too deep, add workers. Twenty workers pull from one queue, each message goes to exactly one of them, throughput scales almost linearly — then the autoscaled workers drain the 9pm spike and scale back down.
On a partitioned log like Kafka, there's one crucial ceiling. Consumers join a named consumer group, and the broker assigns each partition to exactly one consumer in the group — that's what preserves per-partition order; two consumers on one partition would race. Twelve partitions feed at most twelve consumers; a thirteenth joins and sits idle. Partition count is your parallelism ceiling, chosen at design time — so provision with headroom (3–5x expected consumers), because repartitioning later is disruptive. And beware the skewed key: one account producing 40% of all messages makes its partition a hot partition one consumer must drain alone. That's the celebrity problem — Chapter 2.6.
Most components treat a pile-up as an emergency. A queue is the one place in your architecture where a pile-up is the design working. The classic example: marketing wants to push a notification to 10 million users at 9pm sharp. Done synchronously, that's a tidal wave — your notification service and every database it touches must survive an instantaneous 100x spike, meaning 100x hardware that idles 23.9 hours a day.
With a queue: producers dump 10 million messages in a couple of minutes — cheap, since enqueueing is just a fast durable write — and consumers drain at a deliberately steady 5,000 a second. The backlog peaks in the millions and clears in about 33 minutes, while your databases see a flat, boring, survivable load the whole time. The queue converted a spike into a plateau. That shock absorber is the most common reason a queue box appears on a whiteboard: every bursty write path — notifications, feed fan-out, log ingestion, webhook delivery — wants one.
But a backlog is only healthy if it drains. Watch two numbers: queue depth (messages waiting) and — more telling — oldest-message age (how long the head of the line has waited). A depth spike that drains by 9:35 is Tuesday. An age that climbs and keeps climbing means consumers are slower than producers permanently — not a spike, a capacity deficit, and no queue size fixes arithmetic. You must add consumers, make them faster, or slow the producers down. Deliberately pushing back on producers is called backpressure — Chapter 2.7. The queue bought time to react, not immunity.
Shopify's platform hosts the internet's most violent traffic spikes: celebrity "drops," where a merchant releases limited stock and a fanbase descends in the same minute, and Black Friday/Cyber Monday, where the platform peaked around $4.2 million in sales per minute in 2023. Their survival strategy is queues at two layers. Inside the request path, Shopify pushes everything non-essential into background jobs — their job infrastructure routinely queues tens of thousands of jobs per second, handling webhooks, emails, fraud analysis, and analytics off the user's clock, so checkout stays fast while the backlog absorbs the surge. And for flash sales, they queue the humans: when a drop would overwhelm inventory, buyers land in a throttled checkout queue — a waiting room — admitted in batches sized to what the inventory database can survive, instead of a hundred thousand simultaneous purchases stampeding the same stock rows. Same principle at both layers: don't provision for the spike — absorb it, and drain at a rate you chose.
Three archetypes cover almost every real choice. Kafka is a log: messages append to a durable, ordered, partitioned log and are kept for a retention window whether or not anyone read them — consumers just move bookmarks, so five teams can read the same stream and you can replay history. RabbitMQ is a smart broker: it routes messages through exchanges to queues, tracks per-message acks, supports priorities, and deletes messages once consumed. SQS is a managed queue: fewer features, effectively unlimited scale, zero servers to operate.
| Kafka | RabbitMQ | SQS | |
|---|---|---|---|
| What it is | Replicated append-only log, split into partitions | Broker with routing (exchanges → queues), per-message acks | Fully managed queue service (AWS) |
| After consumption | Message kept until retention expires — replay and multiple readers are natural | Deleted once acked — gone | Deleted once acked — gone |
| Ordering | Guaranteed per partition | Roughly per queue; weakens with retries and multiple consumers | None on standard queues; FIFO queues order per message group at limited throughput |
| Throughput | Millions of msgs/sec per cluster | Tens of thousands/sec, comfortably | Standard: effectively unlimited; FIFO: capped per group |
| Ops burden | Heaviest — a stateful cluster you run (or pay a managed service to run) | Moderate — a broker to run and tune | None — that's the product |
| Pick it when | Event streaming, several consumers of the same data, replay/audit, analytics pipelines | Task queues with routing logic, priorities, low-latency delivery | You're on AWS and want a work queue that just works — the default until a requirement rules it out |
The honest guidance, which is also the senior interview answer: for background jobs, the boring managed queue is the right default — reach for a log like Kafka when you actually have its requirements: multiple independent consumers of the same events, replayability, or six-figure messages per second. In an interview, say "a queue — SQS or equivalent" and move on to the semantics: acks, retries, DLQ, idempotency. Interviewers grade what you know about duplicates, not logos.
Queues surface in nearly every design question, and the senior signal is never the queue itself — it's volunteering the failure semantics unprompted. A passing answer sounds like: "Everything the response doesn't depend on goes through a queue — delivery is at-least-once, so consumers are idempotent on the message ID; failures retry five times then dead-letter, and I alert on DLQ depth and oldest-message age." Probes to expect:
A junior asks: "Why did we move the confirmation email to a queue, and why is there all this weird duplicate-checking code in the email worker?" Explain both, in five or six sentences.
The user only needs the card charged and the order saved before we can honestly say "confirmed" — the email can arrive a minute later, so making them wait on a slow email API was pure waste, and when that API went down it took checkout down with it. The queue fixes both: we drop a "send email for order X" message into a durable buffer and reply immediately; workers send emails at their own pace, catching up after outages. The catch is at-least-once delivery: if a worker sends the email but crashes before confirming, the broker can't tell whether the work happened, so it plays safe and delivers the message again. Every message can therefore show up twice, and the duplicate-check — "have I already sent the email for order X?" — is what makes the second delivery harmless. That's idempotency, and it's the other half of using a queue at all: the queue guarantees the work is never lost, and our code guarantees it's never done twice.
A ride-hailing app's "book ride" endpoint does all of this before responding: (a) validate pickup/dropoff, (b) find and assign a driver, (c) charge a booking fee, (d) send the rider a confirmation push, (e) update the driver's earnings dashboard, (f) log to the analytics warehouse. Classify each as sync or async using the "does the response depend on it?" test.
Sync: (a) validation — you can't accept an invalid booking; (b) driver assignment — "ride booked" without a driver is a lie the rider discovers in seconds; (c) the charge — confirming a booking that might fail payment is also a lie. Async: (d) the push — the rider is already staring at the confirmation screen; (e) the dashboard; (f) analytics. The interesting borderline is (b): real systems often return "finding your driver…" immediately and assign asynchronously — legitimate, but it changed the product: a pending state now exists in the UI. That's the key insight: async work either must not change the response's meaning, or the response must be redesigned with an honest pending state.
Your video-processing workers pull transcode jobs from a queue with a 60-second visibility timeout. Most jobs take 30 seconds, but 4K videos take 4 minutes. Users report that some 4K videos appear in their library two or three times. Walk the mechanism, then fix it.
Worker A pulls a 4K job; at 60 seconds the timeout expires while A is still happily transcoding; the broker assumes A is dead and redelivers to worker B; two workers finish the same video and both write results — duplicates manufactured by a config value, no crash required. With a max-receive-count policy, the job might even hit the DLQ while succeeding. Fixes: (1) heartbeat-extend the timeout while work is genuinely in progress; (2) raise the timeout above worst-case processing time — simple, but slows recovery from real crashes; (3) make the output write idempotent on video ID — do this anyway, since at-least-once means duplicates were always possible. The senior answer is 1 and 3 together: timeouts for liveness, idempotency for correctness.
Marketing will push a notification to 20 million users at 9pm. Workers each push 200 notifications/second; the downstream push gateway tolerates at most 8,000 requests/second. How many workers, how long to drain, and what do you tell marketing about "9pm sharp"?
The gateway cap is the real constraint: 8,000/s ÷ 200/s per worker = 40 useful workers (a 41st just gets throttled). Drain time: 20,000,000 ÷ 8,000/s = 2,500 seconds ≈ 42 minutes. Honest message to marketing: sends begin at 9:00 and finish by roughly 9:42 — fine for a promo. The queue let you run 40 modest workers instead of a system that blasts 20M pushes instantly, and it protected the gateway from a spike it explicitly can't handle. Senior touch: this blast must not share a queue with time-critical notifications like OTP login codes — a 42-minute backlog in front of an OTP is an outage. Separate queues per priority class.
An engineer proposes: "Payments must be reliable, so our payment-events pipeline will use one FIFO queue with strict global ordering — every event processed exactly in sequence." Traffic is 4,000 events/second across 2 million accounts. What breaks, and what do you propose instead?
Global order means one consumer processing serially — at even 2 ms per event, a 500/second ceiling against 4,000/second of traffic: the backlog grows by 3,500 events every second, forever. It also adds head-of-line risk: one poison event stalls the company's payments. Fix: scope the ordering claim. No rule requires account A's events ordered against account B's — order matters per account. Partition by account_id (or FIFO message groups keyed on it): strict order within an account, parallelism across 2 million accounts. Then check skew: one giant marketplace account producing a big share of events makes its partition a hot lane one consumer drains alone — Chapter 2.6's problem, worth naming unprompted.
11:58 p.m. on a Tuesday. Your API has been boringly healthy for months. Then the dashboard turns red: traffic just jumped 40x in ninety seconds. You brace for the exciting explanation — front page of Reddit, or an attack. The truth is dumber. One of your customers shipped a nightly sync job with a bug: when a request fails, it retries immediately, in a loop, with no delay. At midnight their 200,000 installed apps all woke up, hit one broken endpoint, and started hammering.
Your database saturates. Requests from your other customers start timing out. And here's the cruel part: every timeout looks like a failure to the buggy clients, so they retry harder. The more your system struggles, the more load it receives. Nobody attacked you. A for-loop with no sleep in it took down your platform, and it will keep it down until someone's phone rings.
Now rewind the night and add one small thing: a rule at your front door — "each customer gets at most 100 requests per minute; past that, we answer instantly with 'slow down' and touch nothing expensive." The buggy fleet still goes berserk at midnight, and your system shrugs. Bad requests bounce off the front door in microseconds; every other customer has a normal night. You read about the whole event in a graph the next morning, with coffee, instead of at 3 a.m., with dread.
That rule is a rate limiter, and the front door it lives in is an API gateway. This chapter is about both: the algorithms that decide "yes or no," the etiquette of saying no politely, the surprisingly tricky problem of counting when you have many servers, and the piece of infrastructure that ties it all together.
Engineers new to rate limiting picture hackers. Wrong picture. The traffic that actually melts APIs, day after day, is almost always legitimate credentials doing illegitimate volumes: a retry loop with no delay, a script that polls every 50 milliseconds, a data team "just backfilling one table" against production. Malice is rare; bugs and thoughtlessness are constant.
So limits exist for four reasons, and only one of them is abuse:
A rate limiter is the fuse box in your house. The fuse doesn't exist because your appliances are evil — it exists because one day a heater will short-circuit, and without the fuse, that one bad appliance burns down the whole house. With it, one circuit clicks off and everything else keeps running. Note where the analogy is honest: blowing the fuse is supposed to happen occasionally. A rate limit that no one ever hits is either protecting nothing or set too high to matter.
Fine — you're convinced you need a limit like "100 requests per minute per customer." Now the real question: what exactly does "per minute" mean? It turns out there are five reasonable answers, and the differences between them decide whether your limiter has holes in it.
Imagine each customer has a bucket that holds up to 10 tokens. Tokens drip into the bucket at a steady rate — say 5 per second — and the bucket simply overflows when full; extra tokens are wasted. Every request must take one token to pass. Token available? Allowed. Bucket empty? Rejected. That's the entire algorithm.
Look at what those two numbers give you. The refill rate caps the long-term average: over any long stretch, a client cannot exceed 5 requests per second, because that's all the tokens that exist. The bucket capacity allows short bursts: a client that was quiet for a while has a full bucket and may fire 10 requests back-to-back. That combination — bounded average, bursts forgiven — matches how real clients behave. A mobile app opens and makes 8 calls at once to render its home screen, then goes quiet for a minute. Token bucket says: fine. A hard "5 per second, ever" limit would punish that completely normal burst.
The implementation is delightfully cheap. You do not run a timer dripping tokens into millions of buckets. You store two numbers per customer — current token count and the timestamp of the last update — and do the refill lazily, only when a request arrives:
elapsed = now - last_update
tokens = min(capacity, tokens + elapsed * refill_rate)
if tokens >= 1:
tokens -= 1
last_update = now
allow
else:
reject # and tell them when to come back
Two numbers per key, a few arithmetic operations per request. This is why token bucket is the industry default. AWS documents its EC2 API throttling as exactly this — Describe* calls draw from a shared bucket of about 100 tokens refilled at roughly 20 per second. Stripe has published their design too: a token bucket per user, state in Redis, guarding every request — alongside a separate limiter on concurrent requests and load shedders that protect critical traffic when the fleet is stressed. When the two most imitated API companies both reach for the same algorithm, that's a strong prior.
A token bucket is a prepaid arcade card that auto-tops-up: the arcade adds one credit every few seconds, the card holds at most ten. Show up after a quiet afternoon and you can play ten games in a burst; stay all day and you're held to the top-up rate no matter what. The card overflowing at ten is the part people forget — being polite for a week doesn't entitle you to a thousand-request rampage on Friday. Your savings are capped at one bucket.
Token bucket permits bursts. Sometimes that's exactly wrong. Suppose behind your API sits a fragile legacy system, or a third-party partner who contractually accepts 50 requests per second, flat. A 200-request burst — perfectly legal under a token bucket — would knock them over.
Flip the metaphor. Requests pour into a bucket with a small hole in the bottom. The bucket is a queue; the hole leaks requests out at a perfectly constant rate — 50 per second, like a metronome. Arrive faster than the leak and you wait in the bucket; fill the bucket completely and new arrivals are rejected. The output is the point: no matter how spiky the input, what comes out the bottom is a smooth, fixed-rate stream. This is a traffic shaper, not just a limiter — nginx's limit_req module is a leaky bucket under the hood, and network routers have shaped packets this way for decades.
The trade-off is the queue itself. Queued requests sit there aging — you've converted "rejected" into "slow," which is sometimes worse: a user staring at a spinner may prefer an honest, instant no. Choose leaky bucket when the downstream needs smoothness; choose token bucket when you just need to bound a client's appetite.
The version everyone invents first: divide time into fixed windows — 12:00:00 to 12:00:59, 12:01:00 to 12:01:59 — and keep one counter per customer per window. Limit 100 per minute? Increment the counter keyed by (customer, current_minute); reject at 100; counters die with their window. In Redis this is one INCR and one expiry. It's the cheapest possible limiter, and plenty of internal systems run it happily.
But it has a hole exactly where the windows meet. A client sends 100 requests at 12:00:59 — window one says fine, 100 of 100. Two seconds later, at 12:01:01, they send 100 more — window two just reset to zero and says fine again. Result: 200 requests in two seconds, from a client who never technically broke the "100 per minute" rule. A fixed window admits up to double your limit in a burst straddling the boundary — and clients find this seam, sometimes accidentally (cron jobs fire at :00 sharp), sometimes on purpose.
The exact fix: stop bucketing time at all. Keep a log of the timestamp of every request per customer (a Redis sorted set works). On each new request, delete entries older than 60 seconds, count what's left, and compare against the limit. The "window" now slides continuously with the clock — there is no seam, no boundary trick, ever. The count is perfectly correct at every instant.
Now run the numbers, because the numbers kill it. Limit of 1,000 requests per hour, 10 million active keys, 8 bytes per timestamp: up to 10M × 1,000 × 8B = 80 GB of raw timestamps — and sorted-set overhead multiplies that several times, with a multi-step data-structure operation replacing one increment on your hottest path. You're spending gigabytes and latency to distinguish "103 requests this minute" from "97." Rate limits are guardrails, not billing — that precision buys nothing. Keep the log for the rare case where exactness is the point (a hard legal or contractual cap). For everything else, there's a compromise so good it became the default at the biggest edge network on the internet.
Keep fixed windows — but keep two counters per key: the current window's and the previous window's. To estimate the count in the true sliding window, take the current counter plus a weighted slice of the previous one, weighted by how much of the previous window still overlaps. Say the limit is 100/minute and you're 15 seconds into the current minute: the sliding window covers the last 45 seconds of the previous minute, so estimate = current + previous × 0.75. The assumption baked in: last window's requests were spread evenly across it. That's false in detail, roughly true in aggregate — and it completely closes the boundary hole, because the previous window's burst still counts against you for a while after the seam.
Cost: two small integers per key, one increment per request. Accuracy: approximate — but how approximate?
When Cloudflare built rate limiting as a product, the constraint was brutal: it had to work for millions of customer domains, enforced across every edge server they run, without melting memory. A timestamp log per visitor per domain was arithmetic suicide at that scale, so they went with the sliding window counter — two counters per key and the weighted-overlap estimate. Then they did the responsible thing: measured how wrong the approximation actually is. They analyzed 400 million requests from 270,000 distinct sources and found the estimator made the wrong call on just 0.003% of requests, with the estimated rate differing from the true rate by about 6% on average — and none of the wrongly-allowed traffic got through at anywhere near abusive volume. In their write-up they note the pleasant irony: the approximation's smoothing across window boundaries isn't a bug, it's the feature — it's exactly what defeats the boundary-burst trick. That analysis is why "sliding window counter" is the answer most senior engineers give by default: near-exact behavior at fixed-window prices, validated at internet scale.
| Algorithm | Allows bursts? | State per key | Accuracy | Reach for it when |
|---|---|---|---|---|
| Token bucket | Yes, up to capacity | 2 numbers | Exact (for its own rules) | Default for client-facing API limits; bursty clients are normal |
| Leaky bucket | No — output is smoothed | Queue (or 2 numbers) | Exact | Downstream demands a steady rate; shaping, not just limiting |
| Fixed window | Up to 2x at the seam | 1 counter | Leaky at boundaries | Rough internal guardrails where 2x briefly is fine |
| Sliding window log | No | One timestamp per request | Perfect | Small scale, or exactness is legally the point |
| Sliding window counter | Slight, bounded | 2 counters | ~exact (0.003% error at Cloudflare) | High-scale general-purpose limiting |
An algorithm needs a key: limit 100 per minute per what? The candidates: API key (the natural unit for a paid API — it maps to the entity you made promises to), user ID (for logged-in product traffic), IP address (the only handle you have before authentication), and endpoint (a search query can cost 100x a profile fetch; expensive routes deserve their own, tighter budgets).
Each key alone has a famous failure. Key only on IP and you punish the innocent: thousands of users behind one university NAT or one mobile carrier's gateway share an address, while a real attacker rotates through cheap proxies. Key only on user and your login endpoint is defenseless — the attacker doing credential stuffing isn't logged in yet, so there's no user to count against. Key only on API key with one global number and a customer's report-generation batch job starves their own production traffic.
So mature APIs layer limits: a generous per-IP ceiling as a crude pre-auth safety net; a per-key limit as the fairness contract; tighter per-key-per-endpoint limits on expensive routes; and often a cap on concurrent in-flight requests, because 50 slow queries running at once can hurt more than 500 fast ones spread over a minute. GitHub's public API shows the layering in the wild: 60 requests per hour unauthenticated (keyed by IP), 5,000 per hour authenticated (keyed by account), and — added later, after watching real abuse — "secondary" limits that trigger on rapid-fire bursts, concurrency, and expensive mutations even when the hourly budget still has room. The evolution is the lesson: they started with one simple number and reality kept adding layers.
A limiter that silently drops requests or returns a vague 500 creates a support ticket and a confused, retrying client. The polite refusal is standardized and you should treat it as part of your API contract:
HTTP/1.1 429 Too Many Requests
retry-after: 30
x-ratelimit-limit: 5000
x-ratelimit-remaining: 0
x-ratelimit-reset: 1755088800
429 says "your fault, and specifically: too fast" — distinct from 403 (forbidden) and 500 (our fault). retry-after tells the client exactly how long to wait; it converts guesswork into instruction. The x-ratelimit-* family — used by GitHub, Stripe, and nearly every serious API — rides along on successful responses too, telling clients their budget, what's left, and when it resets. That last part is quietly the most valuable: well-built clients watch remaining and slow themselves down before hitting the wall. The cheapest 429 is the one you never have to send.
Now the client side, because you will write clients too, and bad clients are how this chapter's cold open happened. The naive retry — fail, retry immediately — is a tiny denial-of-service attack. The textbook fix is exponential backoff: wait 1 second, then 2, 4, 8, capped somewhere sane, so a struggling server sees pressure fall off exponentially.
Except there's a trap hiding in the textbook version. Picture 10,000 clients that all failed at the same instant — which is exactly what happens when a server blips, because the failure itself synchronizes them. With pure exponential backoff they all retry at t=1s. Together. Then all at t=3s. Together. You've turned a crowd into a marching army: waves of 10,000 simultaneous requests in lockstep, each wave as punishing as the original spike. The server gets re-DDoSed by its own well-behaved clients.
The fix is jitter — randomness. Instead of sleeping exactly base × 2^attempt, sleep a random amount: random(0, min(cap, base × 2^attempt)), the variant known as full jitter. Randomness breaks the synchronization; the army dissolves back into a crowd, arrivals spread evenly instead of clumping. AWS's architecture blog ran the simulations that made this famous: full jitter finished the same workload with dramatically fewer total calls and less client waiting than plain exponential backoff. One call to random() is the highest-leverage line in any retry loop.
On September 20, 2015, DynamoDB — the database under a huge slice of AWS itself — went down in us-east-1 for around five hours, and the amplifier was retries. Inside DynamoDB, storage servers periodically fetch their "membership" — the list of partitions they're responsible for — from an internal metadata service. A recently popular feature (global secondary indexes) had quietly grown those membership payloads, pushing fetches close to their timeout. Then a brief network disruption caused a batch of storage servers to miss their renewals. They all re-requested at once; the now-heavier responses overwhelmed the metadata service; servers that couldn't confirm membership took themselves out of service and kept retrying — every failure creating more load on the thing that was failing. The load was so relentless that AWS engineers couldn't even add capacity, and recovery required an almost comic step: deliberately blocking requests to their own metadata service — rate limiting it, after the fact — just so it could breathe while they scaled it up. The collateral damage spread to SQS, CloudWatch, and EC2 auto-scaling. The postmortem's fixes are this chapter in production form: more headroom, stricter limits on how aggressively clients re-request, and honest capacity math on the payload growth. If retries-without-limits can do that to AWS, they can do it to you.
Everything so far assumed someone can count. On one server, trivial — a hashmap. But you don't run one server. You run twenty gateway instances behind a load balancer, and a customer's requests land on all of them. If each instance keeps its own local counter, a "100 per minute" limit quietly becomes "100 per minute per instance" — 2,000 per minute, drifting every time you autoscale. The counter has to be shared.
And shared counting has a classic bug. The obvious code — read the count, check it, write count+1 — is a read-modify-write race. Two instances read "99" at the same moment; both check 99 < 100; both allow; both write 100. Two requests passed through slot number 100 — and under a burst from twenty instances, far more, because the read and the write are separate steps with a gap other servers dive through.
The fix is to make the count-and-decide a single atomic step, done where the data lives. Redis is the standard tool: INCR is atomic — Redis executes commands one at a time, so "increment and return the new value" happens with no gap for anyone to dive through. Each gateway calls INCR and checks the returned value; the race is gone. For token buckets, where the update is "read two numbers, compute a refill, conditionally decrement," you wrap the logic in a small Lua script — Redis runs a script as one uninterruptible unit, giving you an atomic token bucket in five lines. That's Stripe's design, verbatim.
The price is a network hop: every request now pays a round trip to Redis — call it 0.5–1 ms inside a datacenter — and Redis becomes infrastructure you must scale (a single node handles on the order of 100K ops/sec; beyond that you shard counters across nodes by key). For most APIs, 1 ms against a 200 ms budget is noise, and the simple central counter is the right answer. At true edge scale you relax precision instead: count locally, sync in the background, accept a few percent of overshoot.
Here's the judgment call interviewers use to sort seniors from algorithm-memorizers. It's 2 a.m. and your Redis cluster — the one holding all the counters — goes down. Every gateway instance now faces a choice on every request: it cannot count, so does it fail closed (reject everything, stay "safe") or fail open (allow everything, stay up)?
Fail closed sounds cautious and is usually catastrophic: your protection mechanism just converted a Redis outage into a 100% outage of your entire API. The safety system crashed the plane. For general rate limiting, the senior default is fail open, loudly: allow traffic, trip a circuit breaker so you stop paying timeouts on a dead Redis, page someone, and optionally fall back to crude local per-instance limits as a stopgap. The reasoning: the limiter is protection, not correctness — being unprotected for ten minutes is a risk; being down for ten minutes is a certainty.
But say the second half too: some limits must fail closed, because for them the limit is the product. Password-attempt limits (failing open hands attackers a free credential-stuffing window), SMS-sending caps (failing open spends real money), fraud and payout velocity checks. The senior answer is never one policy; it's "fail open for traffic protection, fail closed where the counter is a security or financial control — and we decide per limit, in advance, not at 2 a.m."
Step back and look at what we've accumulated: authenticate the caller, find the right limit, count atomically, reject with proper headers, forward to the right backend. Should every one of your services implement all that? Twenty services, twenty limiter implementations, twenty places to update when the auth scheme changes? No. You put one component in front of everything and give it the job. That's the API gateway: the single front door through which every external request enters your system.
A gateway is a hotel front desk. Guests don't knock on random doors: the desk checks your ID once (authentication), tells you your room (routing), stops the person trying to visit all 400 rooms in an hour (rate limiting), and answers common questions itself (aggregation). Housekeeping and the kitchen — your backend services — get to assume anyone who reaches them was already checked. Where the analogy breaks: a hotel has one desk and survives; you will need several desks with identical behavior, because one desk is a single point of failure.
The gateway's core duties, in the order a request meets them: TLS termination — decrypt HTTPS once at the edge, so backends never touch certificates. Authentication — validate the API key or JWT (a signed token proving who the caller is) once, then stamp the request with a verified identity header that everything downstream simply trusts. Rate limiting — which must come after auth, because the limit is keyed on who you are, and must come before anything expensive, because the whole point is that a 429 costs microseconds, not a database query. Routing — map /orders/* to the orders service, /search to search. And often aggregation: a mobile home screen needs data from four services; rather than make a phone do four round trips over flaky cellular, the gateway fans out over fast datacenter links and returns one response. Plus the unglamorous duty that pays daily: one consistent access log and metrics stream for every request entering the company.
"Everything passes through it" should make you twitch: isn't that a single point of failure? Only if you build it as a single anything. The design above is deliberate: the gateway holds no per-request state — counters live in Redis, sessions in tokens — so it's a fleet of identical, disposable instances behind a load balancer: one dies, nobody notices. The real SPOF risk in a gateway isn't hardware, it's configuration: one bad rule deployed to every instance takes down every route at once — a lesson Cloudflare learned publicly in 2019, when a single bad regex in their web application firewall pinned CPUs across their global edge and took down every site behind them for half an hour. So gateway config gets the most paranoid rollout in your stack: canary to one instance, watch, then spread. And keep the gateway thin — cross-cutting concerns only. The moment product logic ("validate order totals") sneaks in, you're rebuilding the enterprise service bus, a 2000s architecture that died of exactly this: a shared chokepoint that every team must coordinate to change.
One honest preview before we close: everything here counted within one region, one Redis's reach. Serve from twelve regions and "count atomically in one place" collides with the speed of light — a Sydney request can't afford a round trip to a counter in Virginia. The real answers (per-region budgets, async counter merging, bounded overshoot) are a genuinely hard problem with their own chapter (p3c10). Within a region, this chapter is the full playbook.
Rate limiting appears two ways: as a component drop-in during any design ("limiting at the gateway — token bucket per API key, counters in Redis"), and as a full deep-dive ("design a rate limiter"). A passing senior answer commits to an algorithm with a reason, keys it on the right identity, writes the 429 contract unasked, and volunteers the two failure conversations: the read-modify-write race (fixed by atomic INCR/Lua) and fail-open versus fail-closed (judged per limit type). Probes to expect:
A junior asks: "Why would we rate limit our own paying customers? And what actually happens when a request comes in?" Answer in five or six sentences, including how a token bucket decides.
We limit customers because the biggest threat to our API is not hackers — it's our own customers' bugs: one broken retry loop can send us 40x normal traffic and take down the platform for everyone, so a limit turns their bug into their error messages instead of our outage. It's also fairness (capacity is shared, and one greedy script shouldn't starve a thousand polite apps) and cost (every request costs money). When a request arrives, the gateway checks the customer's token bucket: tokens drip in at a steady rate — that caps their long-term average — and the bucket holds a maximum, which is the burst we'll forgive after quiet periods. If a token is there, we take it and the request proceeds; if the bucket is empty, we instantly return 429 with a retry-after header telling them exactly when to come back, without touching the database at all. The counting happens in Redis with atomic operations, because twenty gateway servers incrementing a shared counter the naive way would race each other and let extra requests slip through. And if Redis itself dies, we let traffic through rather than block everything — the limiter is a seatbelt, not the engine.
Estimation drill: you serve 10 million active API keys with a limit of 1,000 requests/hour each. Compare the memory for a sliding window log (Redis sorted set, ~8 bytes per timestamp, assume ~4x structure overhead) versus a sliding window counter. Which do you pick, and what did the numbers just decide for you?
Log worst case: 10M keys × 1,000 timestamps × 8B = 80 GB raw; with ~4x sorted-set overhead, ~320 GB — a multi-node Redis cluster existing purely to remember timestamps. Counter: 2 counters ≈ 16B per key plus key-name overhead (say ~50B total) × 10M ≈ 0.5 GB — one modest Redis node. The numbers just decided the architecture: ~600x memory difference to gain precision you don't need (Cloudflare measured the counter's error at 0.003% of requests). Senior move: state the comparison in one breath and pick the counter — reaching for the log at this scale signals you didn't run the math.
Your fixed-window limiter allows 600 requests/minute. What's the worst burst a client can legally push through in two seconds, and how exactly do they do it? Then explain what a sliding window counter would say to the same client at the same moment.
1,200 requests — double the limit. Send 600 at 12:00:59 (window one: 600/600, legal), then 600 at 12:01:01 (window two just reset: 600/600, legal again). Two seconds, 1,200 requests, zero rejections. The sliding window counter at 12:01:01 sees it differently: estimate = current count + previous window (600) × ~0.98 overlap ≈ 588 before the new burst even starts — the budget is already almost spent, so essentially the entire second burst gets rejected. The weighted tail of the previous window is precisely the memory that fixed windows lack at the seam.
A service goes down for 30 seconds. It has 50,000 clients that poll every 5 seconds and, on failure, retry after exactly 1 second (no backoff, no jitter). Describe the traffic the service faces when it tries to come back up, roughly quantified. Then fix the client policy and describe the new picture.
The outage synchronizes everyone: within seconds, all 50,000 clients are in the 1-second retry loop, so the recovering service faces a wall of ~50,000 requests per second — likely 10-50x its normal load — arriving in synchronized waves. It comes up, gets flattened, falls over, repeats: a metastable failure where the outage sustains itself (this is the DynamoDB-2015 shape). Fix: exponential backoff with full jitter — sleep random(0, min(cap, base × 2^attempt)). Backoff drops the retry volume exponentially over time; jitter is the crucial part here, dissolving the synchronized waves into a smooth spread the service can absorb while recovering. Server side, you'd also shed load early: answer 429/503 with retry-after cheaply at the gateway rather than letting half-finished work saturate the backend.
Sketch the rate limiting for a login endpoint. What do you key on, what algorithm, what numbers, and — the judgment question — does this limiter fail open or fail closed when Redis dies?
You can't key on user ID — the caller isn't authenticated yet. Layer two keys: per-IP (catches single-source floods; keep it generous — carrier NATs put thousands of humans behind one address) and per-attempted-username (catches distributed credential stuffing against one account, which IP limits miss entirely). Token bucket with a tiny capacity — say 5 attempts, refilling 1 per minute per username — because a human's legitimate "burst" is fat-fingering a password a few times, not fifty. Numbers stay small because the cost asymmetry is extreme: a 429 costs you nothing; a compromised account costs everything. And this is the classic fail-closed case: if Redis dies, failing open would hand attackers an uncounted credential-stuffing window exactly when your systems are shaky. Reject (or heavily CAPTCHA) logins during the outage; briefly inconveniencing legitimate users beats silently disabling a security control. Saying "fail open for traffic limits, fail closed here, decided per limit" is the senior answer in one sentence.
What breaks first if your gateway's traffic grows 10x? Walk the request path — load balancer, gateway fleet, Redis counters, backends — and find the weakest link and its fix.
The gateway fleet itself is fine — it's stateless, so 10x traffic means adding instances, which autoscaling does without drama. The load balancer is managed infrastructure with enormous headroom. The weak link is the shared Redis: every request costs it an atomic operation, and a single node saturates around ~100K ops/sec — 10x may blow past that. Worse, sharding counters across Redis nodes by key fixes the aggregate but not the hot key: one giant customer's API key still hammers a single shard (all counting for one key must land on one node — that's what makes it atomic). Fixes, in escalation order: shard counters by key across a Redis cluster; for hot keys, split each counter into N sub-counters incremented round-robin and summed on read; or move to local counting with async sync, accepting a few percent overshoot. Naming the hot-key problem unprompted — not just "add more Redis" — is what the 10x question is fishing for.
Tuesday design review. The new feature is an activity feed. A teammate puts up a slide: "Postgres won't scale for this. I'm proposing MongoDB. Or maybe Cassandra — Netflix uses Cassandra." Heads nod around the room, because nobody wants to be the engineer defending the old database. You glance at the dashboard: the whole product holds 80 GB of data and peaks at 300 requests per second.
Here's the uncomfortable fact hiding under that meeting: 80 GB fits in the RAM of a single rented server, and 300 requests per second is a load Postgres handles while doing a crossword. The proposal isn't engineering — it's fashion. The same scene plays out in interviews every day, and interviewers listen for it, because the database choice is the decision that most cleanly separates "reasons from access patterns and numbers" from "reasons from blog posts."
So treat what follows as a field guide, not a religion: why these databases exist, what each of the four families is actually for, and what's inside the engines — because once you see the machinery, the trade-offs stop being magic and start being arithmetic. And one piece of interview advice will repeat until you're sick of it: pick one database per family and know it deeply. Depth in Postgres plus one NoSQL store beats shallow familiarity with nine logos, every single time.
Relational databases — Postgres, MySQL — are one of the most successful ideas in computing. They give you joins (combine data across tables in one query), transactions (several changes succeed or fail as a unit), and strong consistency (after a write commits, every read sees it). For forty years, that bundle has been the right default.
But the bundle was designed for one machine. Joins and transactions are cheap when all the data sits on one box; the moment your data spreads across fifty machines, a join might mean shipping gigabytes over the network, and a transaction might mean fifty machines agreeing in lockstep — slow, fragile, sometimes impossible under failure.
Amazon's engineers wrote the famous Dynamo paper (2007) after watching their shopping-cart infrastructure strain during holiday peaks. Their observation was the founding insight of the whole movement: the cart didn't need joins, and it didn't need perfect consistency — it needed to never refuse a write, even while servers were failing mid-Christmas. So they built a store that dropped joins and transactions entirely and bought, with that sacrifice, always-on availability and effortless scale-out across cheap machines.
That's what NoSQL is: a trade, not an upgrade. Every NoSQL database removes features from the relational bundle to buy something specific — horizontal scale, write throughput, flexible schema, or one access pattern made absurdly fast. Which means the only sane way to pick one is to know what you're buying and what you're handing over. Three honest reasons to leave Postgres:
"SQL doesn't scale" is not on the list, because it's false at the sizes most systems ever reach. Keep that sentence loaded for design reviews.
Every mainstream NoSQL store belongs to one of four families, and each family is a bet on one access pattern. Learn the mental models, one exemplar each, and the when-to-pick — that's the whole landscape.
Four ways to store your stuff. A cloakroom: hand over your coat, get ticket 47; with the ticket, retrieval is instant — but ask the attendant for "all black coats size M" and you'll get a blank stare. That's key-value. A filing cabinet: everything about the Sharma case lives in one folder — contracts, photos, notes — grab the folder, you have it all; but finding "every case mentioning witness Mehta" means opening every folder. That's document. A diary with one page per day: appending today's entry is instant and entries stay time-ordered, but you must know the date to find anything. That's wide-column. A family tree on the wall: following lines from person to person is what the drawing is for — questions like "how are these two related?" answer themselves. That's graph. None of these is "better than a library with a catalog" (the relational database). They're each better at exactly one kind of question, and worse at the rest.
Mental model: a hash map the size of a building. You store a value (bytes, JSON, whatever) under a key; you get it back by that exact key. No queries, no joins, no "find all values where…". Because the database promises so little, it can be preposterously fast and scale almost without thought — hash the key, that's the machine it lives on.
Exemplars: Redis (in-memory, sub-millisecond, rich value types like counters and sorted sets — you met it in the caching chapter) and DynamoDB (Amazon's managed descendant of the Dynamo paper: on-disk, replicated, single-digit-millisecond reads whether you store one gigabyte or one petabyte).
Pick it for: session storage ("session token → user context"), shopping carts ("user ID → cart"), feature flags, rate-limit counters, cache layers. The tell is in the sentence: if you catch yourself saying "given X, fetch its Y" and never "find all the Ys where…", you're describing a key-value problem.
Mental model: a filing cabinet of self-contained JSON objects. A product document holds its own variants, images, and reviews, nested inside; one read fetches the whole thing — data that would be five joined tables in Postgres. Documents in the same collection don't have to share a shape, so adding a field next sprint touches no schema.
Exemplar: MongoDB. Pick the family when your unit of reading and writing is naturally a whole nested object (product page, user profile, CMS article, insurance claim) and the shape genuinely evolves. The cost shows up the day your access patterns stop matching the folder boundaries: "every document mentioning X" or data that's actually relational — orders referencing products referencing inventory — turns into either duplicated data or joins reimplemented, badly, in application code. Honest footnote: Postgres's JSONB gives you queryable, indexable documents inside a relational database, so "we need nested JSON" alone no longer justifies a second database.
Mental model: a map of maps, and the keys are the entire design. The partition key decides which machines own the data — hash it, land on a node, exactly like the consistent hashing you learned in the sharding chapter. Inside a partition, rows are stored physically sorted by the clustering key. So "all messages in channel 42 between 9:00 and 9:05" is: jump to one machine, one sequential sweep through sorted rows. No scanning, no scatter.
Exemplars: Cassandra and ScyllaDB (same data model; Scylla is a C++ rewrite of Cassandra's Java). Both are leaderless — every node accepts writes, with the quorum machinery from the consistency chapter — so there's no single primary to bottleneck, and adding nodes adds capacity nearly linearly. Pick the family for write-heavy, time-ordered workloads: chat messages, activity feeds, sensor readings, event logs. The cost: you must know your queries when you design the table, because the table is the query. Ask anything the keys weren't designed for and the database shrugs.
In November 2011, Netflix was mid-migration from its own datacenters to AWS and needed to know whether Cassandra could really absorb their write load in the cloud. So they ran a now-famous benchmark: a Cassandra cluster spread across three AWS availability zones, tested at 48, 96, 144, and 288 instances. Throughput scaled almost perfectly linearly — double the machines, double the writes — topping out at over 1.1 million client writes per second on the 288-node cluster, with every write replicated three ways across zones. The whole experiment took about two hours and a few hundred dollars of EC2 time. That linearity is the whole promise of the leaderless wide-column design: no primary to saturate, no re-architecture at each step, just more nodes. It's why Cassandra became Netflix's workhorse for high-volume operational data like viewing history — and why "we'll add nodes" is a credible scaling answer in this family in a way it never is for a single-primary relational database.
Mental model: nodes and edges stored so that walking from a node to its neighbors is a pointer hop, not an index lookup. Exemplar: Neo4j. The honest use case is variable-depth traversal — queries where you don't know how many hops you'll need: "is this new account connected, through any chain of shared devices, cards, or addresses, to a known fraud ring?" In SQL, unknown-depth traversal becomes recursive joins that get uglier and slower with every hop; in a graph store it's the native operation.
Now the honesty the sales deck omits: most "graph problems" are two joins. "Friends of friends" is a self-join, twice, on a follows table — Postgres eats it with an index. "People who bought this also bought" is a two-hop query with a fixed depth — same. In an interview, reaching for a graph database on fixed-depth questions signals you don't know what joins can do. The graph family earns its keep at hop three and beyond, when depth is unknown, or when the traversal itself is the product (fraud detection, dependency analysis, knowledge graphs).
You can use databases for years treating the storage engine as magic. But one distinction explains so many observed behaviors — why Cassandra swallows write floods, why its reads sometimes wobble, what "compaction" alerts mean — that it's worth ten minutes.
Postgres and MySQL organize data in B-trees: a sorted tree of fixed-size pages on disk. To update a row, the engine finds its page and modifies it in place — and also updates every index that references it. Reads are wonderful: any row is a few page hops away. Writes pay the price: finding and rewriting pages scattered across the disk, once per index.
LSM trees (log-structured merge trees — the engine inside Cassandra, ScyllaDB, RocksDB, and LevelDB) refuse to modify anything in place. A write goes two places: appended to a WAL (write-ahead log — an append-only file on disk, so the write survives a crash) and inserted into the memtable (a sorted structure in RAM). Then the database says "done." That's the entire write path — one sequential append plus one memory insert, which is why it's fast: no seeking, no page rewriting, no index shuffling. When the memtable fills, it's flushed to disk as an SSTable — a sorted, immutable file. Updates don't touch old data; they just write a newer version. Deletes write a tombstone — a little marker meaning "pretend this is gone."
The bill arrives on the read path. A row's latest version could be in the memtable or any of a dozen SSTables, so a read may consult several files (helped by per-file Bloom filters — probability tricks that say "definitely not in this file"). Left alone, files pile up and reads keep slowing. So a background process called compaction continuously merges SSTables: it takes several sorted files, merge-sorts them into one, keeps only each row's newest version, drops tombstoned data for real, and deletes the inputs. Compaction is why LSM reads stay acceptable — and it costs real disk I/O and CPU, forever, in the background.
A B-tree is a librarian who reshelves every returned book immediately: returns are slow (walk to the right shelf, every time), but finding any book is instant. An LSM tree is a returns bin by the door: dropping a book takes one second no matter how busy the library gets — but to find a book you must check the shelves and rummage through every bin, newest first. So each night, staff merge-sort the bins onto the shelves and throw out superseded editions. That nightly merge is compaction. And the classic LSM failure is visible in the analogy: if returns flood in faster than the night staff can sort, bins multiply, and every lookup gets slower — the library is drowning not in books but in unsorted bins.
Discord's message store is the best public tour of this whole chapter. They launched in 2015 on MongoDB — right choice for shipping fast. By November of that year they had 100 million messages, the data and indexes no longer fit in RAM, and latency turned unpredictable. So they migrated to Cassandra, modeling messages with a partition key of (channel_id, bucket) — a bucket being roughly ten days of one channel's messages, precisely the bounded-partition thinking above. That design carried them from billions to trillions of messages on a cluster that grew to 177 nodes. Then the engine itself became the pain: a giant channel going viral created hot partitions — traffic hammering the few replicas owning that partition, latency cascading cluster-wide (adding nodes can't fix this; the partition still lives on the same replicas). Cassandra's JVM added garbage-collection pauses — the runtime freezing to reclaim memory — that they tuned endlessly. And compaction fell behind under load, so reads touched ever more SSTables. In 2022 they migrated to ScyllaDB — same data model, C++ engine, no GC, one shard per CPU core — and put Rust "data services" in front that coalesce identical concurrent reads: a thousand clients requesting the same hot message become one database query. Results: 177 nodes down to 72, and worst-case p99 message reads from 40–125 ms to about 15 ms. Notice what survived all three migrations: the access pattern and the partition-key design. The engines were replaceable; the data model was the real asset.
In Postgres, when you want to query by a new field, you type CREATE INDEX and go get chai. It's the single habit that transfers worst to distributed databases — and interviewers love probing exactly this.
A secondary index is an index on anything other than the primary key. On one machine it's cheap. But in a store that's sharded by partition key, data for "status = shipped" lives scattered across every node. Distributed databases handle this in one of two unpleasant ways. Cassandra-style local indexes: each node indexes only its own data, so a query by that field must fan out to every node and merge results — called scatter-gather, with latency set by the slowest node; fine at 5 nodes, grim at 100. DynamoDB-style global secondary indexes (GSIs): the database maintains what is honestly a second copy of your table, repartitioned by the new key and updated asynchronously — so it costs extra storage, extra write capacity on every write, and reads from it are eventually consistent (may briefly lag). Nothing here resembles "type one line, go get chai." In distributed stores, every query pattern must be paid for explicitly — with a duplicate of the data organized for that query.
Follow that principle to its logical end and you get DynamoDB single-table design: list every access pattern up front, then cram all your entity types — users, orders, order items — into one table, overloading generic partition and sort keys (USER#123 / ORDER#2026-08-13#456) so that each access pattern is one cheap key lookup, and related records land adjacent for one-query retrieval. It reads as gibberish next to a relational schema, and that's the point: it's a table organized by query, not by entity. The trade is stark — spectacular performance and scale for the queries you predicted, real pain for the ones you didn't. Which makes the deeper lesson of this whole chapter explicit: relational databases let you defer knowing your queries; NoSQL makes you pay for that knowledge up front.
One more shelf in the field guide. A class of databases — Google's Spanner, its open-source cousin CockroachDB — refuses the founding trade of NoSQL: they offer SQL, real transactions, and strong consistency and horizontal scale-out across machines and regions. Spanner does it with TrueTime, a fleet of GPS receivers and atomic clocks in Google's datacenters that lets nodes bound clock uncertainty tightly enough to order transactions globally. The catch is that physics still gets paid: cross-region commits wait on replication and clock-uncertainty windows, so writes carry tens of milliseconds of floor latency; the systems are expensive; and the operational surface is serious. When someone in an interview says "I need transactions AND multi-region scale," this class is the honest answer — delivered with its price tag, not as a free lunch.
Now the field guide's actual method. Three questions, asked in this order — most people ask them in reverse, which is how 80 GB products end up on Cassandra.
1. What are the access patterns? Not "what data do I have" — how is it read and written? List the top queries and their rough frequencies. "Given a session token, fetch the user" is key-value. "Append events, read one key's recent window" is wide-column. "Read a whole nested object whose shape keeps changing" is document. "Traverse relationships to unknown depth" is graph. "Many entities, cross-referenced, queried in ways I can't fully predict" — that's relational, and it describes most products.
2. What are the consistency needs? If two records must change together atomically — money moves, seat bookings, inventory — you need transactions, and that vetoes most of the NoSQL landscape no matter how well the access pattern fits. Postgres, or Spanner-class if you truly need multi-region scale too. If stale-by-a-second reads are fine (feeds, likes, sensor dashboards), the eventually consistent families stay on the table.
3. Is the scale real? Be honest. Put numbers on it — the estimation chapter's whole job. And here is the sentence this book will stand behind: under roughly a terabyte of data and a few thousand queries per second, the answer is Postgres. One primary, read replicas, a cache in front — that architecture serves companies with millions of users. Choosing it isn't timid; in an interview it reads as senior, because you attached numbers: "At 200 GB and 800 QPS peak this is comfortably Postgres. The signal that changes my mind is sustained write throughput beyond what one primary plus batching can absorb — then I'd move the write-heavy table, and only it, to Cassandra." Commit, justify, name the escalation trigger.
| Family | Mental model | Know deeply | Pick when | It punishes you when |
|---|---|---|---|---|
| Key-value | Giant dictionary | Redis or DynamoDB | Get/put by exact key at high volume: sessions, carts, counters | You need queries, not lookups |
| Document | Folders of JSON | MongoDB | Whole nested objects, genuinely evolving shape | Data is relational; queries cross folder lines |
| Wide-column | Two-level map: partition key → sorted rows | Cassandra/ScyllaDB | Write-heavy, time-ordered, known queries: feeds, messages, telemetry | New query patterns arrive after the table is designed |
| Graph | Walkable relationships | Neo4j | Unknown-depth traversals are the product | Your "graph problem" was two joins |
| Relational | The flexible default | Postgres | Under ~1 TB, mixed/unpredictable queries, transactions | Sustained write volume beyond one primary |
And the interview advice, one more time, because it decides real loops: pick one per family and know it deeply. When you say "Cassandra" in a design, the follow-up will be "what's your partition key?" — and the candidate who answers with (channel_id, day_bucket) and an explanation of why the bucket bounds partition size beats the candidate who knows nine database names and no keys. Depth is checkable in one follow-up question; breadth is not a senior signal at all.
Database choice comes up in every single design round, and it's graded on reasoning, not the logo. A passing senior answer names the access pattern, commits to a store, states the cost, and gives the escalation trigger — four sentences. Probes to expect:
Explain to a junior, in five or six sentences, why Cassandra can absorb writes so much faster than Postgres — and what it gives up in exchange.
Postgres stores rows in sorted structures on disk and updates them in place — every write means finding the right page, rewriting it, and updating every index, which is a lot of disk work. Cassandra never modifies anything in place: a write just gets appended to a log file (so it survives a crash) and dropped into a sorted table in memory, and that's the whole job — sequential and fast. Memory tables get flushed to disk as immutable sorted files, and a background process called compaction keeps merging those files so they don't pile up. The price appears on reads: the latest version of a row might be in memory or in any of several files, so reads do more work and can wobble when compaction falls behind. Cassandra also drops joins and transactions — you have to organize tables around the exact queries you'll run, chosen in advance. So it's a trade: spectacular write throughput and easy scale-out, paid for with read complexity and query flexibility.
Pick a family (and one exemplar) for each, with a one-line justification: (a) session store for 40M daily users; (b) product catalog — nested variants, shape changes weekly, 60 GB; (c) IoT platform ingesting 400K sensor readings/sec, dashboards read one device's last hour; (d) "people who bought X also bought Y" recommendations, exactly two hops.
(a) Key-value — Redis (or DynamoDB if you need durability without managing failover): token → session, tens of millions of lookups, zero queries. (b) Trick question — 60 GB with nested evolving shapes is Postgres with JSONB; document-shaped data is not, by itself, a reason to run a second database. MongoDB becomes defensible if the whole product is document-shaped and the team owns it deeply. (c) Wide-column — Cassandra/ScyllaDB: relentless append-heavy writes, reads are one partition's time range; partition key (device_id, day), clustering by timestamp. (d) Postgres — fixed two-hop traversal is two self-joins on an orders table; a graph database at fixed depth is the over-engineering tell. Neo4j enters when depth becomes unbounded (fraud chains, dependency closure).
You're designing chat storage on Cassandra. A teammate proposes partition key channel_id, clustering key timestamp. What are the two failure modes, and what did Discord actually do?
Failure one: unbounded partitions. A busy channel accumulates messages forever, and its partition grows without limit — multi-gigabyte partitions wreck compaction, repair, and read latency. Failure two: permanent hot partitions. A channel's entire history and all its traffic live on the same few replicas forever; one viral channel melts exactly those nodes, and adding nodes doesn't help because the partition doesn't move. Discord's fix: partition key (channel_id, bucket) where a bucket is about ten days — bounding every partition's size and spreading one channel's load across many partitions over time. General rule worth memorizing: every partition key needs a story for its maximum size and its peak traffic; if either is unbounded, add a dimension (time bucket, hash suffix) to the key.
Your team runs 200 GB in Postgres at 800 QPS peak (85% reads). A proposal argues for migrating to DynamoDB "so we never have to worry about scale again." Write the three-sentence senior response.
Something like: "At 200 GB and 800 QPS, Postgres with two read replicas and a cache is bored — we're at maybe 5% of what this architecture handles, so a migration buys us nothing users can feel and costs us a quarter of roadmap plus every ad-hoc query we run today. DynamoDB would also force us to enumerate every access pattern up front, and our analytics queries alone disqualify that. The trigger for revisiting: sustained write growth toward a few thousand QPS on the hot tables or the working set outgrowing RAM — and then we'd move the specific hot table, not the whole database." Numbers, cost of the migration itself, a named trigger — that's the shape.
What breaks at 10x: your Cassandra cluster ingests clickstream events at 100K writes/sec, p99 read latency 12 ms, disks at 40%. Traffic grows 10x over six months. Name the first three things to go wrong, in order.
First, compaction falls behind: at 1M writes/sec the background merging can't keep pace with SSTable creation, files-per-read climbs, and read p99 degrades from 12 ms toward 100+ ms — writes still look fine, which is what makes it sneaky; watch pending-compactions and SSTables-per-read metrics. Second, hot partitions surface: whatever skew exists in your partition key (one huge customer, one popular page) is now 10x hotter, and those replicas spike while cluster averages look healthy — averages lie, per-partition metrics don't. Third, disk pressure: 10x ingest eats the 60% headroom fast, and compaction needs scratch space (roughly the size of what it's merging), so trouble arrives well before 100%. The senior habit: on LSM stores, monitor compaction debt and partition skew as first-class signals, not just CPU and disk.
Monday morning at your ticket-resale marketplace, and there are three bugs in the queue. Bug one: two customers opened support tickets with the same order number. You sharded the orders database last month — chapter 1.9's medicine — and now shard A and shard B are both happily counting 1, 2, 3 from their own auto-increment columns. Order 4,000,317 exists twice, and they're different orders.
Bug two: search is timing out. Someone typed "taylor swift" into the search box, and the query — a LIKE '%taylor swift%' against ten million listings — took 40 seconds and pinned a database CPU. Fifty people searching at once during a ticket drop turned the whole database into a space heater.
Bug three is the scary one. Your nightly refund job runs on a server. Last week you added a second server for redundancy. Both servers ran the job. A customer got refunded twice, and finance wants a word.
Three unrelated-looking bugs, one shared diagnosis: these are all things a single server gave you for free. One machine hands out numbers in order, greps a small table fast enough, and runs a cron job exactly once. The moment you have many machines, all three guarantees quietly evaporate — and you need to rebuild each one deliberately. Unique IDs, search, and locks: the three utility belts that show up in almost every design interview, usually as a two-minute detour that separates people who've built systems from people who've read about them. Let's earn all three.
On a single database, AUTO_INCREMENT (MySQL) or SERIAL (Postgres) is lovely. One counter, one machine, every insert gets the next integer. It's small (8 bytes), it's sortable, and new rows always land at the "end" of the primary-key index, which keeps inserts fast. Then you distribute, and three separate problems appear.
So distributed systems need IDs that are unique without asking anyone. The obvious first answer is randomness.
A UUIDv4 is 128 bits, 122 of them random — the string that looks like f47ac10b-58cc-4372-a567-0e02b2c3d479. Any machine can mint one with zero coordination, and collisions are a non-issue in practice: you'd have to generate a billion UUIDs per second for about 85 years before you hit even a coin-flip chance of one collision. Uniqueness solved. Two new problems bought.
First, size: 128 bits versus 64, and often stored as a 36-character string, which bloats every index and every foreign key that references it. Second — the one that bites at scale — index locality. Databases keep primary keys in a B-tree, which is a sorted structure. Sequential IDs always insert at the rightmost edge: the same few pages stay hot in memory, and writes are cheap appends. Random UUIDs insert everywhere: every write lands on a random page of a huge tree, forcing the database to drag cold pages from disk, split full pages, and dirty pages all over the file. Insert throughput sags as the table grows, and nobody can point at a slow query — the whole table just got heavier.
A B-tree primary key is a filing room with sorted cabinets. Sequential IDs are a clerk who only ever adds files to the last drawer — the drawer's already open, the stack is right there, it's one motion. Random UUIDs are a clerk who must file each new document in a different random cabinet across the room: walk over, unlock it, shuffle contents to make space, sometimes split a stuffed drawer into two. Same number of documents, dramatically more work per document. The analogy breaks in one way worth knowing: the real clerk gets tired, but the real database gets slower and writes more to disk — page splits mean write amplification, not just latency.
The fix is almost embarrassingly direct: make the front of the ID a timestamp and keep the randomness at the back. That's UUIDv7, standardized in RFC 9562 in 2024 — a 48-bit Unix millisecond timestamp followed by random bits. IDs minted around the same time now sort near each other, so inserts land near the right edge of the B-tree again, like the sequential clerk. Still zero coordination, still effectively collision-proof. If you just need "a good ID" in 2026 and don't control every generator, UUIDv7 is the boring, correct default. Its costs: still 128 bits, and the timestamp prefix means an ID reveals roughly when the row was created — usually fine, occasionally a privacy consideration.
But there's a design that gets you time-ordering, uniqueness, and a compact 64 bits — and it comes with the best origin story in this chapter.
In 2010, Twitter was moving tweet storage off MySQL and onto Cassandra, and hit a wall: Cassandra, being a distributed database, has no auto-increment — there is no single counter to increment. Twitter needed tens of thousands of new tweet IDs per second, generated by many machines with no coordination, in under 2 milliseconds. And there was one requirement stricter than uniqueness: tweets had to sort by time. Timelines are "order by ID descending" — if IDs aren't roughly time-ordered, every timeline everywhere breaks. Fully random IDs were unusable. So they built a small service called Snowflake that packs a millisecond timestamp, a machine number, and a per-millisecond counter into one 64-bit integer — unique with zero runtime coordination, and sortable because the timestamp occupies the most significant bits. They open-sourced it, and the layout became the industry's default answer: Discord uses Snowflake IDs for every message and server, Instagram built a variant (coming up), and Sony published their own "Sonyflake." One honest caveat Twitter documented: the IDs are roughly sorted — two tweets in the same second may be slightly out of order across machines — which was fine, because timelines only need time-ordering at human resolution.
Here's the layout, and — more importantly — the reasoning behind every number:
ORDER BY id is a free time sort and B-tree inserts stay right-edge-hot.One failure mode a senior candidate names unprompted: the clock. The whole scheme trusts the machine's clock to move forward. If NTP (the network time-sync protocol) yanks a machine's clock backwards, the generator could re-mint timestamps it already used — and with the same worker ID and sequence, that's a duplicate ID. Real implementations detect time going backwards and refuse to generate until the clock catches up. Say that sentence in an interview and the interviewer relaxes: you've operated things.
In 2011 Instagram had a famously tiny engineering team and sharded Postgres. They needed sortable 64-bit IDs, looked at Snowflake, and balked — not at the design, but at the operations: running ZooKeeper plus a fleet of ID servers is real on-call surface for a team of a few people. So they folded the same idea into the database: a small PL/pgSQL function on each shard builds the ID at insert time — 41 bits of milliseconds, 13 bits of shard number, and 10 bits from a per-shard counter modulo 1024. Same bit-budget philosophy, zero new infrastructure, and the shard ID is embedded in every primary key, so given any photo ID you can compute which shard holds it. The lesson generalizes beautifully: the best version of a pattern is the one your team can operate. Snowflake-as-a-service and Snowflake-in-a-stored-procedure are the same idea wearing different operational costs.
Flickr hit this problem before Twitter did: they federated their MySQL database into many shards in the mid-2000s, needed globally unique photo IDs, and wrote the solution up in 2010. They considered GUIDs (UUIDs) and rejected them for exactly the index-locality and size reasons above — random 128-bit keys index badly in MySQL. Their solution was almost comically low-tech: a dedicated MySQL box whose only job is a one-row table with an auto-increment column. Issuing an ID is a single REPLACE INTO statement — a MySQL quirk that deletes and re-inserts the row, bumping the counter — then read the counter back. That's a "ticket server": the centralized counter, kept, but isolated on hardware whose only job is counting. One box is a single point of failure, so Flickr ran two: one configured to issue only odd numbers, the other only even (MySQL's auto_increment_increment and auto_increment_offset settings), with clients round-robining between them. Either box can die and the other keeps issuing, and they can never collide. The cost they accepted with open eyes: a network round trip per ID, and IDs that are only roughly ordered across the odd/even pair. It carried billions of Flickr photo IDs. Simple, boring, operable — very senior.
Here's the whole decision compressed:
| Scheme | Coordination per ID | Time-sortable | Size | Reach for it when |
|---|---|---|---|---|
| Auto-increment | Central (the DB) | Yes | 8 bytes | One database, no shards — don't over-engineer past this |
| UUIDv4 | None | No | 16 bytes | Uniqueness is all you need and insert volume is modest |
| UUIDv7 | None | Yes (ms prefix) | 16 bytes | The modern default: no infrastructure, B-tree friendly |
| Snowflake-style | Once per machine at boot | Yes (roughly) | 8 bytes | Huge volume, compact sortable IDs, you can run (or embed) the generator |
| Ticket server | Central, every ID | Roughly | 8 bytes | Moderate volume, you want dead-simple and already run MySQL |
IDs sorted. Next bug: the 40-second search query.
Why did LIKE '%taylor swift%' take 40 seconds? Because of what an index fundamentally is. A B-tree index is a sorted list — sorted by the column's value from its first character. Sorted data answers prefix questions fast: "everything starting with tay" is one contiguous range, found by binary search. But a leading wildcard — "taylor swift appearing anywhere in the text" — has no prefix to search by. The sorted order is useless, so the database reads every row and checks each one. Ten million listings at ~100 bytes each is a gigabyte scanned per search. One query: seconds. Fifty concurrent searches during a ticket drop: the database stops being a database.
You cannot index your way out of this with more B-trees. You need a different data structure — and you already own a copy of it, printed on paper.
The index at the back of a textbook. Nobody finds "sharding" in a 900-page book by reading pages 1 through 900 — that's the full table scan. You flip to the back, where someone has already inverted the book: for each important word, a sorted list of the pages it appears on. "Sharding: 212, 340, 585." One lookup, straight to the pages. A search engine's inverted index is exactly this, built for every word automatically: instead of document → words, store word → list of documents. The list attached to each word is called its posting list. The analogy is honest with one addition: a book index only covers chosen terms; a search index covers every token, and it also remembers enough to rank the pages, which the book index never does.
Two details turn this sketch into a real search engine.
Tokenization — deciding what counts as a "word." The text is lowercased and split; then a stemmer folds word forms together ("flights", "flight" → flight) so a search for one finds the other; and ultra-common glue words ("to", "the") may be dropped since their posting lists contain every document and discriminate nothing. Every one of these choices is a product decision wearing an engineering costume: stem too aggressively and "universal" matches "university."
Relevance — posting lists tell you which documents match; users need them ranked. The classic intuition is TF-IDF, and you can carry it entirely without formulas. Two signals: a document that mentions your term more often is probably more about it (term frequency), and a term that appears in few documents is more informative than one that appears everywhere (inverse document frequency — "goa" beats "cheap" because everything on a travel site is "cheap"). BM25, the default scoring function in modern engines, is this intuition plus two adult corrections: repeating a word has diminishing returns (saying "goa" fifty times shouldn't score ten times higher than saying it five times — hello, keyword spammers), and long documents don't win just for containing more words. That's the whole gist; interviewers want the intuition, never the math.
You won't build this by hand — you'll deploy Elasticsearch (or OpenSearch, or a managed equivalent), which wraps the Lucene library's inverted indexes in a distributed, replicated service. And here is the architectural sentence that matters more than any feature: Elasticsearch is a search index over your data, not the home of your data. The system of record — the store you trust, back up, and can rebuild everything else from — stays your real database. Search gets a derived copy, rebuilt from truth whenever needed.
This rule was written in lost data. In 2014 and 2015, Kyle Kingsbury's Jepsen project — a famous test suite that tortures distributed databases with network partitions — showed that Elasticsearch could lose acknowledged writes: during a partition, two halves of a cluster could each elect a leader (a "split brain"), accept conflicting writes, and silently discard one side's data when the cluster healed. Elastic responded with unusual honesty: for years the company maintained a public "Elasticsearch Resiliency" status page listing known data-loss scenarios, and its own documentation advised keeping the authoritative copy of data in another store. Teams that skipped that advice — using ES as the primary store because "the data's already there, why run two databases?" — were the ones filing the data-loss reports. Elasticsearch 7 (2019) shipped a rebuilt cluster-coordination layer that fixed the known split-brain scenarios, and modern ES is far more robust. But the architecture lesson stands on its own: a search index is optimized for fast reads over derived data, not for being the last copy of anything. Design so that losing the index costs you a reindex, never the data.
If the database is the truth and Elasticsearch is a copy, something must ship changes from one to the other. The tempting answer is dual-write: the application writes to Postgres, then writes to Elasticsearch. Resist it. The two writes can't be made atomic — no transaction spans both systems — so the app crashing between them, or an ES timeout after a DB commit, leaves the copies silently disagreeing, forever, with no record of the drift. The robust pattern is CDC — change data capture: a small service (Debezium is the standard) tails the database's replication log — the append-only record of every committed change that the database already produces for its own replicas — and publishes each change to a queue, where an indexer consumes it and updates Elasticsearch. The database's own commit log is the source of events, so nothing committed can be missed, and if the indexer dies it resumes from where it stopped.
Notice the label on that pipeline: seconds behind. Between the commit, the log tail, the queue, the indexer, and Elasticsearch's own refresh cycle (by default it makes new documents searchable about once per second), your search index trails reality by a few seconds under normal load — and by minutes if the indexer falls behind during a spike. Senior candidates volunteer this instead of hiding it: "search will be eventually consistent, seconds behind the database; a just-posted listing may not be searchable instantly, and I'll re-check anything critical — like 'is this actually still for sale' — against Postgres at purchase time, not against the index." Staleness admitted and bounded is a design; staleness discovered by the interviewer is a hole.
One neighbor to wave at: the search box that completes "tay" into "taylor swift tickets" as you type is not this machinery — typeahead has a 50ms budget and prefix-shaped queries, and gets its own purpose-built design in chapter 3.2. Full-text search and autocomplete are cousins, not twins.
Back to the double refund. Two servers both ran the nightly job because nothing told them not to — each one checked "is it 2am? yes" and went to work. On one machine, you'd grab a mutex, a lock inside the process. Across machines there's no shared memory to put a mutex in, so the lock has to live somewhere all the workers can see: a distributed lock.
The everyday pattern uses Redis, and it's one command:
SET lock:refund-job "worker-A-8f3d2c" NX PX 600000
Unpack it. NX means "only set if the key does Not eXist" — so if two workers race, exactly one SET succeeds, and Redis being single-threaded per key is what makes the check-and-set atomic. That winner holds the lock; the loser sees a failure and walks away. PX 600000 attaches a 10-minute expiry (TTL — time to live), and this part is not optional: if the winner crashes mid-job, the key must eventually vanish on its own, or the lock is held by a ghost forever and no refund job ever runs again. The value — a random token unique to this worker — matters too: when releasing, the worker must delete the key only if it still contains its own token (a tiny Lua script makes the check-and-delete atomic). Delete blindly, and a slow worker can release a lock that has already expired and been acquired by someone else — freeing a lock it no longer owns.
A hotel-desk key for the only meeting room. Take the key, and everyone else who asks is told "occupied" — that's NX. The desk has a rule: keys not returned by checkout time are assumed lost, and the desk cuts a new key for the next guest — that's the TTL, and it's what keeps a guest who left town with the key in their pocket from sealing the room forever. But spot the flaw in the rule: it assumes nobody who overstays is still inside. If your meeting runs long and the desk hands a new key to the next group, two groups now believe they have the room. The desk never checks the room — it only tracks time. That flaw is exactly the bug that follows, and it isn't fixable by choosing a better checkout time.
The TTL creates the lock's fundamental dilemma. Too short, and it expires while the work is still running. Too long, and a crashed holder blocks everyone for the whole duration. And "long enough" doesn't exist, because of pauses you don't control. The canonical villain: a garbage-collection pause. In many language runtimes (the JVM most famously), the runtime occasionally freezes the entire program — every thread — to reclaim memory. Usually milliseconds. Occasionally, on a bad day with a big heap: tens of seconds. The program does not know it was frozen. It experiences no gap. And clocks elsewhere kept moving.
Walk the timeline. Worker A acquires the lock with a 30-second TTL and starts issuing refunds. At second 12, a GC pause freezes it for 25 seconds. At second 30, Redis — dutifully, correctly — expires the key. At second 31, worker B acquires the lock, checks the batch, starts refunding. At second 37, worker A thaws and continues exactly where it left off, believing completely that it holds the lock. Two workers, same refunds, and the double-refund bug is back — this time with a lock, which is somehow more embarrassing. GC is only the most famous freezer: page faults, virtual-machine migrations, an overloaded hypervisor, or a network hiccup between worker and Redis all produce the same shape. The uncomfortable truth generalizes: a lock with a timeout cannot, by itself, guarantee mutual exclusion, because the holder can never be sure in real time that it still holds it.
You might diagnose a different weakness first — "the lock lives in one Redis; what if Redis dies?" — and reach for Redlock, an algorithm by Redis's creator that acquires the lock on a majority of five independent Redis nodes, tolerating node failures. In 2016, Martin Kleppmann (of Designing Data-Intensive Applications) published a widely-read critique, antirez published a sharp rebuttal, and the exchange became the field's favorite argument. Kleppmann's core point survives contact with both essays, and it's the distinction to carry out of this chapter: decide whether your lock is for efficiency or for correctness. An efficiency lock prevents wasted work — two workers computing the same expensive report; if it rarely fails, you pay some duplicate compute and nobody is harmed, and a single Redis SET NX PX is the right amount of machinery (Redlock's five nodes buy little here). A correctness lock prevents disaster — double refunds, two nodes writing the same file. And for that job, Redlock still isn't enough, because more Redis nodes do nothing about the GC-pause timeline above: Redlock's safety quietly depends on assumptions about clocks and pause times that real machines violate. Kleppmann's verdict, roughly: fine for efficiency (though overkill), insufficient for correctness.
So what does protect correctness? A preview, because chapter 2.3 finishes this properly: a fencing token. The lock service hands every holder a number that only increases — A gets token 33, and after A's lease expires, B gets 34. Every write carries its token, and the storage being written rejects any token older than the newest it has seen. When zombie A thaws and writes with 33 after B wrote with 34, storage refuses. Notice what moved: the enforcement left the lock and moved into the resource — the one place a frozen process can't lie to. Getting tokens that are guaranteed to only increase requires consensus (ZooKeeper, etcd), which is chapter 2.3's whole subject. Until then, the operational rule of thumb: single-Redis locks for efficiency, and for correctness make the operation idempotent — safe to run twice, chapter 2.2 — so that when the lock fails you anyway, nothing burns.
These three tools are rarely a whole interview — they're the two-minute probes inside every design, and each has a trap the interviewer is watching for. IDs: the trap is hand-waving "I'll use UUIDs" without knowing the index cost, or naming Snowflake without the bit layout. Search: the trap is treating Elasticsearch as a magic database and never mentioning how it stays in sync. Locks: the trap is believing the TTL solves everything. Probes you should expect verbatim:
A junior asks: "We sharded the database and now IDs collide. Why not just keep one auto-increment table somewhere central?" Answer in five or six sentences, including what you'd use instead and why it's sortable.
A central counter works, but now every insert in the whole company waits on a network round trip to one box — it's a bottleneck and a single point of failure in your hottest path, and sequential numbers also let outsiders count our orders. The trick big systems use is to make IDs unique without asking anyone: pack a millisecond timestamp, a machine number, and a small per-millisecond counter into one 64-bit integer — Twitter's Snowflake layout. The timestamp lives in the high bits, so bigger ID means later ID, and sorting by ID sorts by time for free — which also keeps database inserts fast, because new keys always land at the end of the index. Each machine only coordinates once, at startup, to claim its machine number; after that it can mint about four million IDs per second locally. The one thing to respect is the clock: if it ever jumps backwards, stop generating until it catches up, or you can mint a duplicate. If we don't want to run any of this ourselves, UUIDv7 — a timestamp-fronted random ID — gets us most of the same benefits with zero infrastructure, at twice the byte size.
Bit-budget drill. Design a Snowflake-style 64-bit ID for a logging system: up to 4,000 collector machines, each bursting to 1 million events/second, and the scheme must last 40+ years. Allocate the bits and check each constraint.
Work each requirement into bits. 4,000 machines needs 12 worker bits (2^12 = 4,096). 1M events/sec/machine = 1,000 per millisecond, so 10 sequence bits (1,024/ms ≈ 1.02M/sec) — barely; take 11 (2,048/ms ≈ 2M/sec) for burst headroom. That leaves 64 − 1 (sign) − 12 − 11 = 40 timestamp bits: 2^40 ms ≈ 35 years. Fails the 40-year constraint — and that's the real lesson: the bit budget is a negotiation. Either take the 10-bit sequence and let bursts briefly spin-wait into the next millisecond (41 bits ≈ 69 years, constraint met), or drop to 11 worker bits if you can cap the fleet at 2,048 machines. Any allocation is fine; knowing that the three fields trade against each other, and saying which constraint you relaxed and why, is the senior answer.
What breaks and why. A team uses UUIDv4 primary keys in a MySQL table that's grown to 500M rows. Inserts have gradually slowed and disk I/O keeps climbing, though traffic is flat. Explain the mechanism and give two fixes with their costs.
Mechanism: random keys insert at uniformly random positions in the primary-key B-tree. At 500M rows the tree is far bigger than RAM, so nearly every insert touches a cold page (a disk read), and full pages split (extra writes). Flat traffic, growing tree → each insert costs more I/O over time. Fixes: (1) switch new IDs to UUIDv7 or Snowflake-style so inserts return to the warm right edge — cheap to adopt, but the existing random keys keep the tree fragmented until data ages out or is migrated; (2) migrate the primary key entirely — the clean end state, but rewriting the clustered index of a 500M-row table is a serious online-migration project (chapter 2.11). Bonus senior point: keep the UUID as a secondary unique column for external use and cluster on an internal sequential key — many teams' actual landing spot.
The staleness negotiation. Your PM says: "A sold-out event must NEVER appear as buyable in search. Zero tolerance." Search runs on Elasticsearch fed by CDC with ~5 seconds of lag. What do you build, and what do you tell the PM?
Separate the two promises hiding in the sentence. Promise one — nobody ever completes a purchase of a sold-out ticket: build that with certainty, by re-checking availability against Postgres (the source of truth) at add-to-cart and again inside the purchase transaction. The index is never consulted for correctness. Promise two — sold-out events never even render in results: that's the 5-second CDC window, and true zero would mean re-checking every displayed result against the database on every search, turning your cheap index reads back into the database load ES existed to absorb. Tell the PM: "Purchases: guaranteed, always. Display: honestly a few seconds behind — same as every large marketplace — and I can shrink the window for the hot events with a targeted cache-check on the top results if it matters." Distinguishing display staleness from transactional correctness is exactly the judgment this question probes.
Lock post-mortem. A worker takes a Redis lock (TTL 30s) to process payout batches; batches occasionally take 3 minutes; last night two workers processed the same batch. Your teammate proposes "raise the TTL to 10 minutes." Write the post-mortem's 'fix' section properly.
The 10-minute TTL fixes the observed incident (work outliving the lease) and creates a new one: a worker that crashes mid-batch now blocks all payouts for up to 10 minutes. Better layers: (1) a heartbeat — the worker extends the TTL every 10 seconds while alive, so the lease tracks liveness rather than a guess about duration; a crashed worker stops renewing and the lock frees in seconds. (2) Accept that even heartbeats lose to a long GC pause — the extension arrives after expiry — so make the payout idempotent: each payout row carries the batch ID with a unique constraint, so a second worker's inserts are rejected by the database itself. That last layer is the only one that actually prevents double payouts; the lock's honest job description shrinks to "make duplicates rare," while idempotency makes them harmless. TTL tuning alone is the junior answer; naming the layer that holds when the lock lies is the senior one.
LIKE '%x%' forces a full scan because B-trees only answer prefix questions; an inverted index (term → posting list) makes search a few list intersections, ranked by the TF-IDF/BM25 intuition: rare terms count more, repetition has diminishing returns.SET NX PX with a random token is the standard lock, but a TTL plus a GC pause can produce two holders who both believe they're alone — so locks with timeouts can't alone guarantee mutual exclusion.Friday evening, flash sale. Your company split its monolith last year, and everyone's proud of the clean new architecture: an order service, a payment service, an inventory service, each with its own database. A customer taps "Buy" on the last discounted PS5. The order service creates the order. The payment service charges her card — ₹41,999, gone from her account. Then the order service calls the inventory service to reserve the console… and the call times out, because inventory is mid-deploy.
So now the universe contains: one charged card, zero reserved consoles, and an order row that says CREATED. The customer refreshes. Nothing. She tweets a screenshot of her bank statement. Support escalates. An engineer — you — gets to spend Saturday writing a script that finds every order in this half-finished state and refunds it by hand.
Here's the bitter part: two years ago, in the monolith, this bug was impossible. Order, charge, and stock lived in one database, wrapped in one transaction. Either all three writes happened or none did. You didn't build that guarantee — the database gave it to you for free, and when you split the data, you gave it back without noticing.
This chapter is about how to get it back — or more honestly, about the three patterns for living without it, and how to talk about them the way a senior engineer does. This topic is quietly one of the highest-signal areas in senior interviews, because it has a wrong answer that sounds right (2PC), a right answer that sounds messy (sagas), and a trap hiding in almost every whiteboard diagram (the dual write).
In the monolith, the checkout was one block:
BEGIN;
INSERT INTO orders ...; -- create the order
INSERT INTO charges ...; -- record the payment
UPDATE stock SET qty = qty - 1 WHERE item = 'ps5';
COMMIT;
That all-or-nothing property — atomicity, the A in ACID — is a promise a single database makes about its own rows. The moment orders, charges, and stock live in three different databases owned by three different services, no single BEGIN … COMMIT can span them. A distributed transaction — one business operation that must update data in multiple places — stops being a statement and becomes a sequence of network calls. And every gap between two calls is a place where the process can die, the network can drop, or the next service can be down.
Put numbers on it, because numbers are what make this real in an interview. Say each of the three steps succeeds 99.9% of the time — a perfectly respectable service. The chance that all three succeed is 0.999³ ≈ 99.7%. So about 0.3% of checkouts hit at least one failure mid-sequence. At 100,000 orders a day, that's 300 half-finished orders, every single day. Partial failure isn't an edge case you'll get to later. It's a scheduled daily event, and your design either handles it automatically or an engineer handles it by hand on Saturdays.
So the question becomes: how do you make a multi-service operation behave all-or-nothing-ish? Computer science has a famous answer from the 1970s. It's worth understanding well — mostly so you can explain why you're not going to use it.
Two-phase commit (2PC) makes several databases commit together by adding a boss. One node plays coordinator; the databases doing the writes are participants. The protocol has two rounds:
2PC is a wedding ceremony. The officiant (coordinator) asks each party in turn: "Do you take…?" — that's the prepare phase, and "I do" is a binding yes vote. Only after hearing every "I do" does the officiant declare "I now pronounce you…" — the commit. Now imagine the officiant faints right after the "I do"s and before the pronouncement. The couple can't leave — they promised. They can't marry — nobody with authority said so. They just… stand there, frozen at the altar, until the officiant wakes up. That frozen state is exactly what happens to databases in 2PC, except the couple is holding row locks that every other request in your system is queued behind.
On paper, 2PC gives you real atomicity across machines. In practice, across microservices, it fails four separate audits:
There's a practical audit too: 2PC across heterogeneous systems requires everyone to speak a common protocol (the traditional standard is called XA), and most of the things in a modern design simply don't. Kafka doesn't. Your payment provider's HTTP API certainly doesn't — you cannot 2PC a call to a payment gateway. The pattern quietly assumes a world of cooperating enterprise databases that mostly no longer describes production systems.
One honest footnote, worth one sentence in an interview and no more: 2PC is alive and well inside single distributed databases, where the vendor controls everything. Google Spanner runs 2PC across shards — but each participant and the coordinator state are themselves replicated with Paxos, so no single machine death can strand the protocol. That's the fix for 2PC's blocking problem, and it costs an engineering effort roughly the size of Spanner. Between your microservices, you don't have that. So at the whiteboard, the senior line is: "I'm not using 2PC across services — it holds locks while blocked on a coordinator, and a coordinator crash leaves participants in doubt. I'll use a saga." Two sentences, then move on.
A saga gives up on one big atomic transaction and replaces it with a sequence of small local transactions — each one inside a single service, each one committing immediately and releasing its locks. In exchange for a plan B: every step ships with a compensating transaction, a new forward action that semantically undoes it. If step 4 fails, you run the compensations for steps 3, 2, 1 — in reverse — and the system ends up in a state that means "this never happened."
Note the word semantically. A compensation is not a database rollback — the original transaction committed ages ago; there's nothing to roll back. You can't un-charge a card; you issue a refund. You can't un-reserve stock by time travel; you add the unit back. The undo is a business action, visible in the real world (the customer may see charge-then-refund on her statement), and it's your job to design one for every step.
| Step | Local transaction | Compensation |
|---|---|---|
| 1. Order service | Create order, status PENDING | Mark order CANCELLED |
| 2. Payment service | Charge the card | Refund the charge |
| 3. Inventory service | Reserve one unit | Release the reservation |
| 4. Order service | Mark order CONFIRMED | — (last step; nothing after it can fail) |
Now replay Friday's disaster. Payment succeeds, inventory times out and — after retries — reports failure. The saga runs compensations backwards: refund the charge, cancel the order. The customer sees "order failed, payment refunded" instead of silence. No Saturday script. The 300 daily half-finished orders now clean themselves up.
Two design rules fall out immediately. First, not everything can be compensated. You can't un-send an email, un-ship a package, or un-dispense cash. So order your steps to put hard-to-undo actions last — after the last step that can fail. The step after which you can no longer turn back has a name, the pivot: before it, failure means compensate backwards; after it, failure means push forwards (retry until done). Second, know when to retry and when to compensate. A timeout or a 500 is transient — retry it. "Card declined" or "out of stock" is a business fact — no number of retries will change it; compensate. Conflating those two is a classic mid-level tell.
The name, by the way, is older than microservices: Hector Garcia-Molina and Kenneth Salem coined "sagas" in a 1987 paper about breaking up long-running transactions inside a single database. The microservices world dug the idea out of the library because it was suddenly the exact shape of the problem.
Somebody has to know what step comes next and what to do on failure. There are exactly two answers.
Choreography: no boss. Each service listens for events and reacts. The order service emits OrderPlaced; the payment service hears it, charges, emits PaymentCaptured; inventory hears that, reserves, emits StockReserved; the order service hears that and confirms. Failures are events too: PaymentFailed triggers whoever needs to compensate. The saga's logic lives nowhere — it emerges from everyone's subscriptions.
Orchestration: a boss. One component — the orchestrator, usually living in the service that owns the business flow — holds the state machine. It commands: charge this card; awaits the reply; commands: reserve this unit; and on failure it walks its own list of completed steps and commands the compensations. Every order's exact state lives in one queryable place.
Orchestration is a wedding planner: she calls the florist, then the caterer, then the band, checks each one off, and if the venue cancels she personally phones everyone to unwind the bookings. Choreography is a potluck: there's no planner, everyone just knows "when I see the invite, I bring a dish." The potluck needs no manager and scales beautifully — but when the dessert person silently flakes, nobody notices until dinner, and finding out who was supposed to react to what means interviewing every guest. That's the trade, honestly: choreography removes the manager and with it the single place you can ask "where is order 123 stuck?"
| Choreography (events) | Orchestration (coordinator) | |
|---|---|---|
| Coupling | Loose — services only know event names | Orchestrator knows every step |
| "Where is order 123?" | Reconstruct from event logs across services | One row in the orchestrator's state table |
| Adding a step | Touch several services' subscriptions | Edit one state machine |
| Timeouts / stuck flows | Every service handles its own; easy to miss | Orchestrator owns all timers and escalation |
| Failure risk | Invisible flow, cyclic event spaghetti | Orchestrator outage stalls new sagas; can grow into a god-object |
| Sweet spot | 2–3 steps, fire-and-forget reactions | Money flows, 4+ steps, anything support must inspect |
The senior move is to commit with a reason, not to present the table and shrug. A good verbatim line: "This flow moves money across four services, and support will need to see exactly where any order is stuck — I'll orchestrate it from the order service. For one-way reactions like 'send a confirmation email when an order confirms,' choreography: just subscribe to the event, nothing to unwind." Both patterns, each where it belongs, decision anchored to visibility and rollback needs.
By the mid-2010s, Uber had the saga problem at company scale: business flows like driver onboarding — document checks, background verification, activation — span many services and can take days, and dozens of teams were each hand-rolling the same fragile machinery of queues, cron jobs, and "current_step" database columns to track where each flow stood. Every homegrown orchestrator had the same bugs: state lost on crashes, flows stuck with no owner, retries and timers reinvented badly. Maxim Fateev and Samar Abbas — who had met at Amazon around 2009 building Simple Workflow Service, AWS's coordination engine — reunited at Uber's Seattle office in 2015 and built Cadence: a shared orchestration engine that records every step of a workflow as an event history in a database, so when a worker crashes, the workflow replays its history and resumes exactly where it left off, timers and retries included. Orchestration code went from infrastructure every team rebuilt to a platform they all shared; Uber open-sourced it in 2017 and runs a huge number of internal flows on it. In 2019 the pair left to found Temporal — a fork of Cadence — and today companies like Netflix, Datadog, and Coinbase run their sagas on it. The lesson for your whiteboard: orchestration-with-durable-state is so consistently needed that it became a product category ("durable execution"). Mentioning that category is staff-level garnish; knowing why the orchestrator must persist its state — because the orchestrator can crash mid-saga too — is senior table stakes.
ACID transactions gave you a second gift you may not have noticed losing: isolation, the I — nobody sees a transaction's changes until it commits. Sagas have no isolation at all. Each local transaction commits immediately and is instantly visible to the whole world, while the saga may still fail and compensate. Other requests can observe — and act on — a state that is about to be undone. It's the distributed cousin of a dirty read.
Concretely: between "charge card" and "reserve stock," another customer queries availability and gets told the unit is gone (it's reserved for a payment that will fail — she walks away from a sale you actually could have made). Or a "total revenue today" dashboard counts a charge that gets refunded two seconds later. Or, nastier, another workflow makes a decision — extends credit, triggers a shipment — based on an order that then gets cancelled.
The everyday countermeasure is one you should volunteer at the whiteboard: the semantic lock, better known as a pending state. Don't write final values mid-saga; write in-progress markers. The order is PENDING, not CONFIRMED. The stock isn't decremented; a reservation row exists with an expiry. Every reader is expected to treat pending as "in flight": the UI shows "processing…", the dashboard excludes pending charges, and a competing checkout can decide whether reserved-but-unpaid stock counts as available. You've rebuilt a small, explicit, application-level version of the isolation the database used to give you for free. The full taxonomy of countermeasures (commutative updates, version checks, re-reads before the pivot) is staff-level garnish; pending states are the 80% answer and are expected at L5.
Now shrink the problem, because the deadliest version of it hides in a single service. Nearly every event-driven design contains this innocent pair of lines:
db.commit(order) // write to MY database
kafka.publish("OrderPlaced", ...) // tell the world
Two systems. No transaction can span them — your database and your message broker cannot commit together. This is the dual-write problem, and both possible orderings are broken:
And no, wrapping it in try/catch doesn't fix it (a crash executes no catch blocks), and retrying doesn't fix it (the process that knew to retry is the thing that died). Interviewers plant this trap deliberately: they wait for you to draw an arrow to the database and another to Kafka, then ask what happens if you crash between them. Finding it in your own diagram before they ask is one of the cleanest senior signals this topic offers.
The fix is almost disappointingly small, which is why interviewers love hearing it by name. You cannot atomically write to two systems — so don't. Write both things to one system, the one that already gives you atomicity: your own database.
Alongside your business tables, add an outbox table. When you save the order, insert the event as a row in the outbox — in the same local ACID transaction:
BEGIN;
INSERT INTO orders (...);
INSERT INTO outbox (event_type, payload, created_at)
VALUES ('OrderPlaced', '{...}', now());
COMMIT;
Now it's physically impossible for the order to exist without its event, or the event without its order — one transaction, one database, the A in ACID doing what it's always done. A separate process, the relay, reads unsent outbox rows and publishes them to the broker, marking each row sent. If the relay crashes mid-publish, it restarts and re-reads unsent rows. Nothing is ever lost; the event sits safely in the outbox until it's delivered.
The relay comes in two flavors. The simple one polls: SELECT * FROM outbox WHERE sent = false every couple of hundred milliseconds. The elegant one uses change data capture (CDC): instead of querying tables, it tails the database's own write-ahead log — the append-only journal every database keeps of committed changes — and turns each committed outbox insert into a published message. Debezium is the standard open-source tool for this. CDC adds no query load and reacts in near real time; polling is an afternoon of work. Either is a correct answer; naming the trade-off is the senior version.
One price tag, and you must volunteer it: the outbox delivers at-least-once. If the relay publishes an event and crashes before marking the row sent, it will publish it again on restart. Duplicates are not a bug in your implementation — they are a guaranteed, designed-in behavior, and every consumer must handle them: processing the same event twice must have the same effect as processing it once. That property is idempotency, it's the outbox's inseparable partner, and it's the entire subject of the next chapter.
As Airbnb broke its Rails monolith (internally nicknamed "monorail") into services through the mid-2010s, dozens of downstream systems needed to know the moment core data changed — a reservation created, a listing price updated — to refresh search indexes, caches, risk checks, and dependent services. Teams gluing this together with dual writes and periodic re-syncs got exactly the drift this chapter predicts: consumers whose view of the data quietly disagreed with the source database. Airbnb's answer was SpinalTap, a change-data-capture service they open-sourced in 2018: it tails MySQL's binlog — the database's own journal of committed changes — plus DynamoDB streams, and publishes every committed mutation to Kafka as a standardized event, in commit order. That last phrase is the whole point: the events aren't what application code claims happened, they're what the database actually committed, so publish-without-write and write-without-publish become structurally impossible. It's the outbox idea graduated into platform infrastructure: stop trusting application code to remember to tell the world; read the truth from the database's log and tell the world yourself.
First, the rule. Memorize it in this exact shape, because it converts this entire chapter into three sentences you can deliver at the whiteboard:
Single-service DB write that must also emit an event → transactional outbox. Multi-service business operation that needs rollback → saga (orchestrated, if money or support visibility is involved). 2PC → name it, then explain in two sentences why you're not using it.
Second, calibration — because the failure mode at senior isn't ignorance of these patterns, it's misjudging how much of this to volunteer, in which order, unprompted:
| L5 — volunteer this unprompted | Staff-level garnish — one sentence, only if the flow invites it |
|---|---|
| Spot the dual write in your own diagram the moment you draw "save + publish," and fix it with the outbox, by name, with the mechanism (event row in the same local transaction; relay publishes after) | CDC internals — binlog tailing, Debezium, ordering guarantees per partition key |
| Say "at-least-once, so consumers must be idempotent" in the same breath as the outbox | Broker-level exactly-once machinery and its limits |
| Choose saga over 2PC for any multi-service flow, with the two-sentence why-not (blocking; in-doubt participants holding locks when the coordinator dies) | Where 2PC legitimately lives: inside Spanner-class databases, coordinator state itself replicated via Paxos |
| Sketch concrete compensations for order → payment → inventory (refund, release, cancel), and order steps so non-compensatable ones come last | Formal pivot-transaction vocabulary; retriable-vs-compensatable step taxonomy |
| Name the isolation gap — "another request can see the half-done state" — and fix it with pending states / semantic locks | Full countermeasure catalog: commutative updates, version checks, re-reads |
| Persist the orchestrator's own state (it can crash mid-saga too) and give every stuck saga a timeout and an owner | Durable-execution platforms — Temporal/Cadence — and their replay model |
The left column is the pass. The right column, delivered instead of the left column, is a fail — garnish on a missing meal. Delivered after it, in single-sentence portions, it reads as depth.
This topic is rarely its own question. It ambushes you mid-design, the moment your diagram spans two datastores. What a passing answer sounds like: you catch the inconsistency yourself, name the pattern, state its guarantee and its price, and keep moving — thirty seconds, no lecture. The probes to expect:
A junior teammate asks: "Why can't we just use a transaction across the order and payment services, like we do inside one database?" Explain the problem and what we do instead, in five or six sentences.
A transaction is a promise one database makes about its own rows — the moment the data lives in two services with two databases, there's no single thing that can make that promise. There is an old protocol, two-phase commit, that fakes it with a coordinator who asks everyone "ready?" then "commit" — but while a participant waits for the verdict it must hold its locks, so if the coordinator dies, rows stay locked and requests pile up behind them. Instead we use a saga: break the operation into small steps that each commit normally in one service, and give each step an undo — refund the charge, release the stock. If a step fails, we run the undos in reverse, so the system ends up as if the order never happened. The catch is that between steps the world can see the half-done state, so we mark things "pending" until the saga finishes. And wherever we write our database and also announce an event, we write the event into an outbox table in the same transaction and let a relay publish it — so the write and the announcement can never disagree.
Take the four-step order saga (create pending order → charge card → reserve stock → confirm order). For a crash or failure at each of the four points, say what the recovery is — and whether it's a retry or a compensation. Then answer: which failure is the most dangerous if your orchestrator does not persist its state?
Fail at step 1: nothing committed anywhere — nothing to do (the client retries checkout). Fail at step 2, card declined: business failure — compensate step 1 (cancel the pending order). Fail at step 3, out of stock: compensate steps 2 and 1 in reverse — refund, then cancel. Timeout at step 3 with unknown outcome: retry first (timeouts are transient and the reserve call should be idempotent); only after retries exhaust do you compensate. Fail at step 4: everything real already succeeded — push forward, retry the confirm until it lands (past the pivot you never go backwards). Most dangerous without persisted orchestrator state: a crash between step 2 and 3 — the charge exists, but no surviving record says a saga was mid-flight, so nobody ever refunds it. That's exactly why the orchestrator's step log must be in a database, not in memory.
Design the saga for a hotel booking: reserve the room, charge the card, email the confirmation. Put the steps in the right order, name each compensation, and identify the pivot.
Order: reserve room (compensation: release the room) → charge card (compensation: refund) → send email (no compensation exists — you can't unsend it). The email must come last precisely because it can't be compensated: it may only fire once nothing after it can fail. The pivot is the charge: before it completes, any failure walks backwards (release the room); after it, you push forward (retry the email until it sends — and email sending had better be idempotent, or a retry double-sends). Bonus consideration: the room reservation should carry an expiry (a semantic lock with a TTL) so a saga that dies un-noticed doesn't hold the room forever.
A teammate's design doc reads: "On signup we insert the user row, then publish UserCreated to Kafka so the search indexer and the welcome-email service pick it up. Both operations are in a try/catch with three retries, so we're covered." Find the flaw, fix it, and name the new problem your fix introduces.
It's a dual write. The try/catch and retries live in the process that can die — a crash between the insert and the publish executes no catch block and no retry, leaving a user that search and email never hear about (or, if publish goes first, an announcement for a user that doesn't exist). Fix: transactional outbox — insert the user row and a UserCreated event row in the same database transaction; a relay (polling, or CDC via something like Debezium) publishes from the outbox to Kafka and marks rows sent. New problem introduced: at-least-once delivery — if the relay crashes after publishing but before marking, the event goes out twice, so the indexer and email service must be idempotent (dedupe on the event id; "welcome email sent" recorded so a duplicate doesn't send twice).
Estimation drill: your checkout saga has 4 steps, each 99.95% reliable, at 500,000 orders/day. (a) How many orders per day need compensation or recovery? (b) Your refund compensation itself calls a payment provider that's 99.9% reliable — what does that imply about how compensations must be built?
(a) P(all four succeed) = 0.9995⁴ ≈ 0.998, so about 0.2% of orders — roughly 1,000 per day, or one every ~90 seconds — hit a failure mid-saga. At that rate, compensation is not an exception path; it's a production feature that needs its own metrics, dashboard, and alerts. (b) ~1,000 compensations/day × 0.1% failure ≈ one failed refund per day. Compensations must therefore be durable and retried from a persistent queue — never fire-and-forget — with a dead-letter queue and a human runbook for the ones that exhaust retries, because "we dropped a refund" is the incident that ends up on Twitter. A saga is only as trustworthy as its worst compensation.
It's the last Friday of the month — payday. At 9:41pm, a user of your food-delivery app taps Pay ₹499. The spinner spins. Five seconds. Ten. The app gives up: "Something went wrong. Please try again." She sighs, taps Pay again, it works instantly, and the biryani is on its way.
Three days later, support ticket #4187: "You charged me twice."
Here's what actually happened. Her first request reached your server. The charge succeeded. The database committed. Then the response — the little 200 OK — died somewhere in a flaky mobile network. From her phone's point of view, a request that never arrived and a response that never returned look exactly the same: silence. So the app retried. And your server, seeing what looked like a brand-new request, charged her again.
You can't fix this by "making the network reliable" — networks fail; that's physics plus economics. Retries are therefore mandatory: a client that doesn't retry silently drops orders, which is worse. And if retries are mandatory, duplicates are mandatory. The only question left — the question that decides payment-adjacent interviews — is who absorbs the duplicates.
When a request times out, exactly one of three things happened: it never reached the server; it reached the server and the work failed; or the work succeeded and only the response was lost. The client cannot tell which. Not "it's hard to tell" — no protocol, no handshake, no number of acknowledgments lets two machines on an unreliable network be certain they agree on what happened.
This has a famous name: the Two Generals problem. Two generals on opposite hills will win only if they attack at the same time, so they send messengers through an enemy-held valley — and any messenger might be captured. "Attack at dawn." Did it arrive? Send a confirmation back. Did that arrive? Confirm the confirmation… there is no finite number of messengers after which both generals are certain, because the last message sent is always unacknowledged. This is why "exactly-once delivery" is marketing shorthand: over a lossy network it's impossible. What's actually possible is better-defined — we'll get there in a minute.
So you get exactly two honest delivery choices. At-most-once: send, don't retry — sometimes the work happens zero times. At-least-once: retry until acknowledged — sometimes it happens twice. For money, "sometimes zero" means lost orders, so every payment system on earth picks at-least-once — and therefore receives duplicates, on purpose, every day. Duplicates aren't an edge case. They're the contract.
So how does your bank avoid double-charging you every time your phone flakes? The reframe that unlocks everything: you can't have exactly-once delivery, but you don't need it. You need exactly-once effect. Let duplicates arrive freely — and build the receiver so that processing a request twice has the same result as processing it once. That property is idempotency: doing an operation N times leaves the same state as doing it once.
At-least-once delivery + idempotent processing = exactly-once effect. That equation is the real meaning behind every "exactly-once" claim you'll ever read — including Kafka's famous "exactly-once semantics," which is at-least-once delivery plus transactional deduplication under the hood. When an interviewer says "exactly-once," a senior candidate says this sentence back, unprompted.
An elevator call button is idempotent. Press it five times — impatiently, like all of us — and exactly one elevator comes, because the system's state is "a request for this floor exists," and pressing again re-asserts a fact that's already true. A chai vending machine is not: press five times, get five cups, pay five times. "Charge ₹499" is a vending machine; our whole job is turning it into an elevator button. Where the analogy breaks: the button is naturally idempotent because its state is a flag. "Charge the card" isn't naturally anything — idempotency has to be bolted on, and how you bolt it on is exactly what interviews probe.
Some operations are idempotent for free: "set status to CANCELLED." But "add ₹499 to the merchant's balance" is not, and rewriting it as "set balance to ₹52,499" doesn't rescue it — another payment landing in between gets silently erased. Real systems do something more general: they detect the duplicate and refuse to redo the work. For that, the server needs a durable memory of what it has already done — and the request needs a name.
On August 1, 2012, Knight Capital — then one of the biggest stock-trading firms in the US — deployed new order-routing code to its eight production servers. The deploy missed one. Worse, the release reused an old feature flag, and on that eighth server the flag activated dead test code from years earlier called Power Peg — code built to keep sending child orders until a counter said "the parent order is filled, stop." The code that updated that counter had been moved years before, so the eighth server never saw "done." It fired orders continuously into the open market: millions of orders, around four million executions across roughly 150 stocks, in about 45 minutes. The loss was over $440 million — more than the company's cash on hand — and Knight had to be rescued and was ultimately acquired. Strip away the finance and the root cause is this chapter's theme: work executing with no durable record of what was already done. One missing "have I already processed this?" check is a company-ending event.
The standard mechanism is the idempotency key. The idea is small; the correctness details are where interviews are won and lost.
The client names the intent. When the user taps Pay, the client generates a unique key — a UUID — and attaches it to the request, typically as an Idempotency-Key header. The crucial rule: the key is born with the logical operation (this tap, this order's payment), not the HTTP attempt — attempt 1 and attempt 2 carry the same key. That's the entire point: a server-generated key would make every retry look brand new. Client-generated, reused across retries. Non-negotiable.
The server remembers what it did. First time it sees a key, the server does the work and records three things in an idempotency table: the key, a hash of the request body, and the response it produced. When the same key arrives again, the server skips the work and replays the stored response — the retry gets a 200 OK with the same charge ID, never knowing its first attempt succeeded. If a known key arrives with a different request hash, that's not a retry, it's a client bug (often a key reused for a new purchase) — reject it with an error like 422; never replay, never execute.
The key record and the business record must commit in the same ACID transaction. One BEGIN: insert the key row, insert the charge row, write the response, COMMIT. Store them separately and you've re-created the original problem one layer down: crash after the key but before the charge, and every retry replays a success for work that never happened; crash the other way and the retry re-executes — double charge. Atomicity makes the crash cases fall out for free: if the server dies mid-work, neither row committed, so the retry finds no key and executes fresh, correctly. This is also why the idempotency table lives in the same ACID database as the business data — not Redis. SETNX (Redis's set-if-not-exists write) plus a separate database write is two systems pretending to be one transaction; one crash or eviction between them and your dedup is fiction. (Redis as a lookup cache is fine — the source of truth stays transactional.)
The race you must name before the interviewer does: two requests with the same key in flight at once — the client's timeout fired while attempt 1 was still running. Handle it in the database, where atomicity lives: put a UNIQUE constraint on the key column and insert the key row first, marked "in progress," before doing the work. The second request's insert hits the constraint, and you choose a policy: block on the winner's row lock until it commits, then replay its stored response — or fail fast with a 409 "already in progress," which is what Stripe does. Blocking is kinder to clients; fail-fast protects your connection pool from piled-up waiters. Pick one and say why. What you may never do is let both proceed.
Scope and TTL are decisions, not defaults. Scope the key per account and per operation — UNIQUE(account_id, endpoint, key) — so two merchants who generate the same UUID never collide, and a key used on "create charge" can't suppress a "create refund." And keys can't live forever: pick a TTL longer than your longest realistic retry window — including a mobile client that queued the request offline and syncs an hour later. Stripe keeps keys for 24 hours, now the industry default (there's even a draft IETF standard for the header). Too short and a late retry becomes a real double charge; longer costs only storage, and the numbers below say storage is a non-issue.
Stripe made idempotency keys famous by making them boring. Every mutating endpoint in their API accepts an Idempotency-Key header; in 2017 engineer Brandur Leach published "Designing robust and predictable APIs with idempotency," laying out the machinery this chapter teaches. Stripe stores each key with a fingerprint of the request's parameters and the response it produced; a replayed key returns the recorded response — the original outcome, whatever it was — without re-executing. Reuse a key with different parameters and you get an error; send the same key concurrently and the second request gets a 409 rather than a second execution. Keys are pruned after 24 hours. The telling detail: Stripe's client libraries automatically generate keys and reuse them across their built-in retries — so a developer who has never heard the word "idempotency" still can't double-charge a customer over a flaky connection. That's the mature end state: the safety is on by default, and you have to work to turn it off.
Duplicates handled. Now for the second habit that separates people who've run money systems from people who've read about them: how you store money. The CRUD instinct is a balance column — UPDATE accounts SET balance = balance - 499. That design fails slowly, then suddenly. Crash between debiting one account and crediting another, and money vanishes. A bug overwrites a balance and you have no way to know what it should be — the history was overwritten in place. A customer asks "why is my balance ₹3,240?" and your honest answer is "because that's what the column says."
Every bank and every accountant since 15th-century Venice uses the same fix: the double-entry ledger. Three rules. First, it's append-only — rows are added, never updated or deleted. Second, every movement of money is recorded twice: a debit on the account it leaves, a credit on the account it enters, and within one transaction total debits must equal total credits. (Real accountants have a richer debit/credit convention; "debit = out, credit = in" is the engineer's simplification — say you're simplifying.) Third, a mistake is fixed by appending a reversing entry — equal and opposite — never by editing the wrong row, which stays visible forever, followed by its correction.
| txn_id | account | debit (paise) | credit (paise) | memo |
|---|---|---|---|---|
| t_9001 | wallet:user_42 | 49900 | — | order #551 payment |
| t_9001 | clearing:processor | — | 49900 | order #551 payment |
| t_9042 | clearing:processor | 49900 | — | refund of order #551 |
| t_9042 | wallet:user_42 | — | 49900 | refund of order #551 |
Why 49900 and not 499.00? Because money is stored as integers in minor units — paise, cents — never as floats. Floats can't represent most decimal fractions exactly: 0.1 + 0.2 is 0.30000000000000004 in every language using IEEE floats, and that error smeared across ten million daily transactions means books that disagree with the bank's by an amount nobody can explain. Integer paise everywhere, with explicit rounding rules at the few places division happens (fees, taxes, conversion). Say it in the interview, one sentence, unprompted: "amounts are integer minor units, never floats." Tiny sentence, lots of credibility.
Notice what the ledger gives you: a balance is now a derived fact, not a stored one — SUM(credits) - SUM(debits) over an account's entries. At scale you won't run that SUM on every read, so you keep a materialized balance — but the relationship never flips: the cache serves reads, and a reconciliation process continuously re-derives balances from the ledger and screams if they disagree. When they differ, the ledger wins, always — it's the record of what happened; the balance is merely an opinion about it.
A ledger is the shopkeeper's bound notebook, written in pen. Every sale, every payout — one line, never torn out, never erased. Wrote ₹500 instead of ₹50? You don't scribble over it; you add a new line — "correction, minus ₹450" — and both stay. The cash drawer is the "balance," but nobody trusts the drawer: at closing you re-add the notebook column and count the cash, and if they disagree, the notebook is right. A balance column without a ledger is a shop that keeps only the drawer — fast to check, impossible to audit. The analogy is honest on one more point: the notebook works because it's one notebook. Ledgers hate being split casually, which brings us to consistency.
Now the stance interviewers wait to hear, stated as a commitment: there is no eventual consistency on the ledger. Ledger writes go through one ACID primary and commit atomically — every leg of the transaction, plus the idempotency key, in one commit. Plenty of the surrounding system can be eventual: analytics replicas, cached balances, notification fanout. But two entries that must balance cannot be "eventually" balanced — in the gap, money exists in two places or neither, and Part 1 taught you exactly what async replication does during that gap. Scale it the way banks do: shard by account, so single-account transactions stay ACID on one primary; cross-shard transfers become a two-phase or saga-style dance you accept reluctantly and audit heavily. Latency is the price; for money, you pay it and say you're paying it on purpose.
A card payment isn't one event — it's a slow conversation. Created: intent recorded. Authorized: the bank approves and holds the customer's money — nothing has moved yet. Captured: you claim the held money — now the charge is real. Settled: it lands in your bank account, a day or two later, minus fees. Off the happy path: failed (declined, expired) and refunded — which, in ledger terms, is a new pair of reversing entries, not an edit.
You learn most of these transitions from webhooks — HTTP callbacks the processor sends you: "payment.captured," "payment.failed." Webhooks are the same hostile network story, from the other direction. The processor delivers them at-least-once (duplicates), retries for hours or days when your endpoint is down (huge delays), and promises no ordering ("captured" can arrive before "authorized," because a retry of an older event lands after a newer one). A consumer that trusts arrival order or arrival count is a time bomb.
The defense is two rules stacked. Rule one: the consumer is idempotent. Every webhook carries an event ID; record processed IDs — same-transaction trick as before — and a redelivered event becomes a no-op. Rule two: events are proposals, not commands. A webhook doesn't set state; it requests a transition, and the state machine only permits legal moves. When "payment.authorized" arrives while you're already captured, captured→authorized isn't a legal move — the event is recognized as stale and dropped. That's what "tolerate reordering" means mechanically: the state machine, not the arrival sequence, is the source of order. Belt-and-suspenders, straight from Stripe's docs: treat the webhook as a doorbell, not a letter — on receipt, fetch the payment's current state from the processor's API rather than trusting the event's payload snapshot.
Idempotency keys absorb duplicates. The state machine absorbs disorder. Now the uncomfortable senior question: what catches the failures those two can't see? Your webhook endpoint 500s for three days and the processor's retries give up. A consumer bug writes ₹499 as ₹4,990. The processor has an incident and never emits the event at all. Each of these leaves your ledger quietly wrong, and nothing in the request path will ever notice — the request path only checks what arrives, not what should have arrived.
The answer is gloriously boring: reconciliation. Every day, the processor publishes a settlement file — a plain file listing every transaction settled on your behalf: IDs, amounts, fees, dates. A daily batch job matches it line-by-line against your ledger, and every discrepancy lands in a named bucket. In their file, not your ledger: a missed webhook — ingest it, book the entries, alert. In your ledger, not their file: a payment stuck in limbo — investigate before the customer does. Amounts differ: usually fees or currency conversion (book adjustment entries), sometimes a real bug (page someone). The one-liner worth memorizing: webhooks are the fast path, reconciliation is the truth path — reconciliation catches what webhooks miss. Mature shops three-way match — ledger against settlement file against bank statement — so even the processor's own mistakes get caught.
Volunteering reconciliation unprompted is one of the strongest "has actually operated money" signals in an interview — it's the layer that exists only if you've asked, "and what catches everything else?"
Calibration, so you neither undershoot nor gold-plate. At L5, the left column is volunteered, not extracted — if the interviewer has to ask, the signal is already weaker:
| L5 must volunteer | Staff-level garnish (fine to name, don't camp on it) |
|---|---|
| "Retries mean duplicates; exactly-once delivery is impossible — I get exactly-once effect via at-least-once plus idempotent handling" | Two Generals / FLP theory beyond one intuition sentence |
| Full key mechanics: client-generated, reused across retries, stored with request hash + response in the business transaction, unique-constraint race handling, TTL and scope with reasons | Kafka transactional internals; exactly-once inside stream processors |
| Double-entry append-only ledger, balances derived (or cached with reconciliation), integer minor units, reversing entries | Multi-currency ledger normalization, FX revaluation entries |
| "No eventual consistency on the ledger" — one ACID primary per account shard, and why | Hot-account mitigation via sub-account balance buckets (name it if probed at 100x) |
| State machine + idempotent, reorder-tolerant webhook consumer | Formal saga orchestration frameworks across many services |
| Daily reconciliation against the settlement file, with the mismatch buckets | Regulatory reporting, audit-compliance specifics |
And keep the numbers doing the deciding, because this topic tempts people into over-building. Key table at 10 million payments a day, roughly a kilobyte per row: about 10 GB a day, and with a 24-hour TTL a steady state of 10–20 GB — one Postgres table with a nightly delete. Anyone proposing a dedicated cluster for that has stopped doing arithmetic. Reconciliation at the same scale: 10 million rows against 10 million is two sorted files and a merge — minutes on one machine. The senior version of this topic is small numbers, small boxes, and correctness obsession; the junior version is big words and Kafka.
Payment-adjacent questions ("design a wallet," "design checkout," "integrate Stripe") are graded almost entirely on this chapter. A passing senior answer narrates the duplicate's journey unprompted: key born at the client, absorbed transactionally at the server, money appended to a ledger, state machine guarding transitions, recon sweeping behind. Probes to expect verbatim:
A junior asks: "Why can't we just have exactly-once delivery? And if we can't, how does Stripe manage to never double-charge anyone?" Answer in five or six sentences.
When a request times out, the sender can't tell whether the request was lost or the reply was lost — and no amount of extra acknowledgments fixes that, because the last message is always unconfirmed. So a network can only offer "maybe never" or "maybe twice," and for payments we choose "maybe twice" and retry until we get an answer. The trick is making "twice" harmless: the app attaches the same unique key to every retry of one payment, and the server keeps a record of keys it has already processed. If a key comes back, the server doesn't charge again — it resends the answer it saved the first time. The key record and the charge are saved in one atomic transaction, so no crash can leave one without the other. Nobody achieved exactly-once delivery — duplicates arrive all the time — but each is recognized and absorbed, and the card sees exactly one charge.
Size the idempotency table: 20 million payment requests a day, roughly 1 KB per stored row, 24-hour TTL. How big is it, where does it live, and what would have to change before you'd consider anything fancier?
20M × 1 KB ≈ 20 GB per day; with a 24h TTL, a steady state of roughly 20–40 GB. That's one indexed Postgres table — in the same database as the charges, because the design depends on committing the key with the charge — plus a nightly delete (or drop yesterday's time partition, which is cheaper). Nothing fancier is justified until the payment database itself shards; then the key table shards with it by the same account key, so key and charge still share a transaction. If your instinct said DynamoDB or Redis, notice what got lost: atomicity with the business write — the one property that made the mechanism correct.
Write the ledger entries (account, debit, credit, in paise) for: customer pays ₹799 for an order; your platform keeps a ₹79 fee; the restaurant gets the rest. One transaction or several? Check your invariant.
One ledger transaction, three legs, committed atomically: debit wallet:customer 79900; credit merchant:restaurant 72000; credit revenue:platform_fees 7900. Debits 79900 = credits 72000 + 7900 — invariant holds. Splitting into two transactions ("customer pays platform," "platform pays restaurant") also balances, but a crash between them leaves the restaurant unpaid with money parked in a platform account, so you'd need a sweeper to complete the second leg. The one-transaction version has no in-between state — that's why it's the default. Say the trade-off out loud.
What breaks if the server generates the idempotency key when a request arrives, instead of the client generating it?
Everything, quietly. The key's job is to give one logical operation one name across retries — and only the client knows attempt 2 is a retry of attempt 1, because only the client experienced the timeout. A server-generated key names the HTTP attempt, not the intent: each retry gets a fresh key, looks unique, and executes. You've built the full apparatus — table, constraint, replay logic — and it will never once fire. A favorite interview trap, because the broken version contains all the right components wired to the wrong trigger. The key is born where the intent is born.
Webhooks for one payment arrive in this order: payment.refunded, payment.captured, payment.authorized. Your stored state is created. Walk each event through your consumer and give the final state.
Event 1, refunded, arrives in state created: not a legal move. Don't apply it, but don't drop it either — park it (dead-letter or re-check later), and better, fetch the payment from the processor's API: the true state is refunded, so fast-forward through the legal chain (created→authorized→captured→refunded), booking ledger entries in order. Working from events alone: refunded parks, captured arrives and parks (or applies if your machine allows the compound move created→captured), authorized applies, then the parked events replay in legal order. Final state either way: refunded, every transition recorded. Teaching point: arrival order carried no information; the state machine plus an authoritative fetch reconstructed the truth.
Find the bug: "We do SETNX key in Redis; on success we process the payment in Postgres and set a 24h expiry on the key; on failure we return 409." Name two failure sequences that double-charge or wrongly block.
Sequence one (wrongly blocks): SETNX succeeds, the server crashes before the Postgres commit. The charge never happened, but the key exists — every honest retry gets 409 until expiry, and the order is silently lost. Sequence two (double-charge): the charge commits, then Redis restarts without persistence or evicts the key under memory pressure — the retry's SETNX succeeds and the payment executes again. Sequence three (subtle): attempt 2 correctly gets 409 while attempt 1 is in flight — but no stored response exists to replay later, so the client never learns the outcome. One root disease: dedup record and business record in different systems can't commit atomically. The fix is moving the key into the payment database's transaction, with Redis demoted to an optional cache.
Friday evening. Your payment service runs the way the replication chapter taught you: one primary database taking writes, one replica standing by. A small watchdog script pings the primary every two seconds; if it misses five pings in a row, it promotes the replica. You've tested this. It works.
Tonight, a top-of-rack switch starts silently dropping packets. The watchdog can't reach the primary, waits its ten seconds, and promotes the replica. But the primary isn't dead — it's on the wrong side of a sulking switch, happily accepting writes from the half of your app servers that can still reach it. Now two databases both believe they're the boss. Payments land on both sides. Some customers get charged twice, some orders exist in only one place, and Monday will be spent hand-merging transaction logs.
This is split brain, and back in p1c08 you were promised a real fix later. This is later. The fix is one of the deepest ideas in distributed systems — consensus — and here's the good news: at the senior bar you are not expected to implement it. You're expected to do something rarer. Know exactly what it guarantees, know where to rent it, and know the one famous way people who rent it still manage to corrupt their data.
Start with why the watchdog was doomed. It made a decision based on silence: "the primary didn't answer, so it must be dead." But over a network, silence proves nothing. A dead server, a slow server, a dropped packet, and a server frozen mid-garbage-collection all look identical from the outside: no reply. There is no test that distinguishes them within any finite timeout. This isn't an engineering gap you can tune away with better pings — it's a hard fact about distributed systems.
So any single observer that promotes a new primary based on a timeout is guessing. Sometimes the guess is right. When it's wrong, you get two primaries — and notice the watchdog itself is just another process that can die, stall, or get partitioned. Adding a second watchdog to watch the first only doubles the number of guessers.
The escape is to change the question. "Is the primary dead?" is unanswerable. "Do a majority of us agree who the boss is?" is answerable — and it has a property that saves you.
Take a fixed group of servers — say five. Define a quorum as any majority of them: three out of five. Now enforce one rule: nothing counts — no leader is elected, no write is durable — until a quorum has agreed to it.
The magic is a piece of arithmetic so small it's easy to miss: any two majorities of the same group must share at least one member. Three out of five and another three out of five overlap in at least one node, no matter how you pick them. That overlap is the whole trick. Two candidates can't both win an election in the same round, because they'd each need three votes and there are only five voters — some voter would have to vote twice, and we forbid that. And whatever the last quorum agreed to, the next quorum is guaranteed to contain at least one node that remembers it. News can never be lost between decisions.
Your housing society has a five-member committee, and any decision needs three signatures. Two rival factions can't both pass conflicting resolutions in the same meeting — three plus three is six signatures, and there are only five pens in the room, so at least one member would have to sign both, which the rules forbid. And when the committee meets next month with a different mix of three, at least one person in the room was present for last month's decision and can say "we already resolved that." Overlap prevents forked decisions and preserves memory. The analogy breaks in one place: committee members can lie or change their story; quorum systems assume nodes are honest but crash-prone, which is why the bookkeeping must be written down, not remembered.
Now do the sizing math, because interviewers love this and it takes ten seconds. A cluster of 2f+1 nodes tolerates f failures while still forming a majority:
| Cluster size | Quorum | Failures tolerated |
|---|---|---|
| 3 | 2 | 1 |
| 4 | 3 | 1 |
| 5 | 3 | 2 |
| 6 | 4 | 2 |
| 7 | 4 | 3 |
Look at rows 3 and 4. A four-node cluster tolerates exactly as many failures as a three-node cluster — one — but its quorum is bigger (three acks instead of two, so agreement is slower) and it has one more machine that can break. The fourth node buys you nothing and costs you twice. That's why coordination clusters are always odd-sized: 3 for most teams, 5 when you want to survive two failures or do maintenance on one node while still tolerating a crash. Seven is for the deeply paranoid; beyond that, quorum latency eats you.
So majorities give you a safe way to vote. But who runs the vote? Nodes crash mid-election. Votes arrive out of order. A candidate dies after winning but before telling anyone. Handling every one of those interleavings correctly is a genuinely hard algorithm — and that algorithm has a name.
Consensus is the problem of getting a group of unreliable machines to agree on a value — or a sequence of values — such that once agreed, the decision is permanent, even as nodes crash and networks partition. The two famous algorithms are Paxos (Leslie Lamport, published 1998) and Raft (Diego Ongaro and John Ousterhout, 2014). For the interview, you need Raft at concept level — three ideas, no more.
Time is chopped into terms, each with a number that only ever goes up. Each term has at most one leader. Every message a node sends is stamped with its current term, and every node follows one reflex: if you see a higher term than yours, update yourself and, if you were a leader, stop being one. A leader from term 3 who was frozen while the cluster moved on to term 4 will have every one of its messages rejected — the stamps give it away. File that reflex away; it comes back later in this chapter wearing a different hat.
The leader proves it's alive by sending heartbeats. If a follower hears nothing for a randomized timeout (typically 150–300 milliseconds), it assumes the leader is gone, increments the term, and asks everyone for votes. Each node grants at most one vote per term — first come, first served — so getting a majority of votes makes you the one true leader of that term. The randomization matters: it makes it unlikely that two followers time out simultaneously and split the vote forever.
The leader accepts commands, appends them to an ordered log, and sends them to followers. An entry is committed once a majority has written it. Because of quorum overlap, any future leader must have been elected by a majority — which must include at least one node holding every committed entry — and Raft's voting rules ensure the candidate with the most complete log wins. Committed entries survive any election. That's the whole guarantee.
Why did Raft win over Paxos in practice? Not correctness — Paxos is proven. Understandability. Raft's paper is literally titled "In Search of an Understandable Consensus Algorithm"; the authors ran a study on students and showed they could learn and apply Raft far more reliably than Paxos. That sounds like an academic nicety until you remember that someone has to implement, debug, and operate this thing at 3am. An algorithm your engineers can hold in their heads is an algorithm with fewer catastrophic bugs.
Google built Chubby (described in Mike Burrows' 2006 paper) as a company-wide lock service: a five-replica Paxos cluster exposing tiny files and locks. GFS used it to elect its master; Bigtable used it for master election and membership — thousands of Google systems outsourced their "who's the boss?" question to one battle-hardened service. Then the Chubby engineers wrote a second, brutally honest paper, "Paxos Made Live" (2007), about actually implementing Paxos: the algorithm fits on a page, but the production system needed thousands of lines of C++ to handle disk corruption, membership changes, and snapshots — and they found that once you add all that, you've built a system the original proofs no longer cover. Their conclusion, from the team best positioned on Earth to implement consensus: the gap between the paper algorithm and a real fault-tolerant system is enormous. There's a delicious epilogue in Google's SRE book: Chubby's global cell became so reliable that teams started assuming it could never fail and wired it into paths that couldn't survive its absence. Chubby's SREs responded by deliberately taking Chubby down for planned outages, forcing dependent teams to fix their assumptions. When your lock service is so good you must break it on purpose to keep people honest, that's the strongest possible argument for renting one instead of writing one.
You'll meet consensus in two packagings, and it's worth telling them apart out loud in an interview.
Standalone coordination services. ZooKeeper (born at Yahoo, runs a Paxos-family protocol called ZAB), etcd (Raft, the brain of Kubernetes), and Consul (Raft, popular for service discovery). Strip away the branding and each is the same product: a tiny, deliberately slow, extremely honest key-value store with three superpowers — atomic compare-and-swap, watches (get notified when a key changes), and keys that expire when their owner dies. That's the whole product. It's enough to build leader election, service discovery, distributed locks, and config distribution — which is why nearly every serious infrastructure project has one of these underneath.
Consensus embedded inside data systems. CockroachDB and TiKV run a Raft group per shard of data, so every write is replicated by consensus before it's acknowledged. And Kafka spent fifteen years leaning on ZooKeeper before replacing it with KRaft — its own built-in Raft quorum for cluster metadata; Kafka 4.0 (March 2025) removed ZooKeeper support entirely. Notice what Kafka did there: a team of world-class distributed-systems engineers took a multi-year, multi-KIP project to replace one consensus implementation with another. That's the size of this undertaking when professionals do it deliberately.
Which sets up the sentence that should come out of your mouth at the whiteboard: "I'd use etcd for this. The guarantee I'm buying is a linearizable compare-and-swap on a small key-value store, with leases and watches — and I'll hang leader election off that." Naming the tool and the guarantee is the senior move. Proposing to implement Raft inside your service is the wrong answer — not because you couldn't sketch it, but because judgment is the thing being graded, and the judgment here is that consensus is a solved, rentable primitive with a decade of hardening (Jepsen, the industry's distributed-systems torture-testing project, has humbled plenty of homegrown attempts). You don't smelt your own steel to build a shelf.
So how does your payment service actually use one of these stores to pick a boss? With a lease: a lock that expires unless renewed.
In etcd: every candidate tries to create the same key — /election/payments-primary — with "create only if it doesn't exist." One wins. The key is attached to a lease with a TTL (time-to-live), say 10 seconds, and the winner heartbeats a keep-alive to renew it. If the leader dies, renewals stop, the lease expires, the key vanishes, and the watchers — every standby was watching that key — race to create it again. New leader, no human involved, no watchdog guessing.
ZooKeeper does the same dance with ephemeral nodes: keys tied to a client's session that vanish automatically when the session dies. The polished version, ephemeral-sequential nodes, appends an auto-incrementing number to each candidate's node; lowest number leads, and each candidate watches only the node just ahead of it, so a leader's death wakes exactly one successor instead of stampeding the whole herd.
Why a TTL at all? Because of the first section: you can't tell dead from slow. The lease converts that unanswerable question into a contract — "you're the boss for 10 seconds at a time, and silence forfeits the crown." Elegant. Robust. And it contains a trap that has become the classic senior-interview probe, so let's walk into it slowly, eyes open.
Your job scheduler must never run the same job twice, so a worker takes the lock, gets the lease, and starts processing. Somewhere mid-job, the JVM decides it's time for a stop-the-world garbage collection pause. Not a normal one — a bad one. The heap is huge, the pause runs 40 seconds. (Multi-second GC pauses are routine at scale; multi-minute ones are documented. The same freeze also arrives as VM migration stalls, page-in from swap, or a laptop-lid-close in dev.)
During those 40 seconds the worker sends no keep-alives, so from etcd's view it went silent, and the lease correctly expires. Worker 2 acquires the lock — correctly. Worker 2 starts writing — correctly. Then worker 1's GC finishes. Here's the horror: a paused process has no idea it was paused. Its next instruction after the freeze is exactly the one it was about to execute — a write. It checked the lease before the pause; it has no reason to check again. It writes. Two workers are now writing under the same lock, and every guarantee in your design just quietly evaporated. This isn't hypothetical: Martin Kleppmann's write-up of this failure cites a real HBase bug where exactly this sequence — GC pause, expired ZooKeeper lease, awakened client — corrupted files in HDFS.
Sit with the uncomfortable part: no lock service can fix this alone. Longer TTLs just delay failover. Checking the lease "right before" the write doesn't help — the pause can land between the check and the write. The process cannot detect its own freeze before its next instruction. Any protocol that relies on the old holder gracefully stepping aside relies on code that may be frozen.
The fix moves the enforcement to the only party guaranteed to be awake: the thing being written to. It's called a fencing token. Every time the lock is granted, the lock service issues a number that only ever increases — grant 33, then 34, then 35. Every write carries its token, and the storage system keeps the highest token it has seen and rejects anything lower. Now replay the disaster: worker 2 writes with token 34; storage records 34; worker 1 wakes and writes with token 33; storage compares 33 against 34 and refuses. The stale writer is fenced off — hence the name.
Hotels stopped using brass keys for a reason. With a brass key, checkout requires physically retrieving the key — and a guest who wanders off with one can still open the door next week. With keycards, the hotel doesn't chase anyone: issuing a new card for the room silently invalidates every older card, and the enforcement lives in the door, which checks each card against the latest issue. The old guest can wave their card all day; the door doesn't argue, it just doesn't open. Fencing tokens are keycards, and your storage system is the door. The analogy's one gap: hotel doors check "is this the current card," while fencing compares numbers — which also lets storage keep accepting the current holder's writes if a new grant hasn't happened yet.
Two practical notes that read very senior when volunteered. First, you rarely have to build the token counter: ZooKeeper's ephemeral-sequential node number and etcd's mod_revision on the lock key are monotonically increasing integers handed to you at grant time — the fencing token is free; the work is making your storage check it (a token column plus a conditional update usually suffices). Second, look back at Raft's terms. A term number is a fencing token — a monotonic epoch that unmasks stale leaders. It's the same idea at two altitudes, and interviewers light up when a candidate connects them. If your resource is too dumb to check anything — a plain file write, a third-party API — you can't fence it, and the honest answer is to make the operation idempotent (p2c02) so a duplicate is harmless, or route the writes through something that can check.
Sooner or later someone on your team proposes distributed locking with Redis, because Redis is already deployed and SET lock_key holder_id NX PX 30000 is one line. For multi-node safety, Redis documents an algorithm called Redlock: acquire the lock, with TTLs, on a majority of five independent Redis nodes. In 2016 Martin Kleppmann published a famous critique ("How to do distributed locking"), Redis creator Salvatore Sanfilippo (antirez) published a famous rebuttal ("Is Redlock safe?"), and engineers have been throwing the two links at each other ever since. Here's the honest summary you should carry into an interview.
Kleppmann's two objections: Redlock issues no fencing tokens, so the GC-pause disaster above goes undetected; and its safety leans on timing assumptions — bounded clock drift, bounded pauses, bounded network delay — that real systems violate (clocks jump when NTP corrects them; pauses are unbounded, as we just saw). Antirez countered that the timing assumptions are reasonable in practice and that many uses don't need fencing. Both are right, because they're talking about different kinds of locks — and that distinction, not the drama, is the senior takeaway:
| Efficiency lock | Correctness lock | |
|---|---|---|
| Purpose | Avoid duplicate work | Prevent corruption |
| If two holders happen | Wasted CPU, a duplicate email, an annoyed user | Corrupted data, double-charged money, a violated invariant |
| Acceptable tool | Single Redis SET NX PX — Redlock is already overkill | Consensus-backed lock (etcd/ZooKeeper) plus fencing tokens checked at the resource |
| Test to apply | Ask "what exactly happens if two clients hold this at once?" — the answer sorts every lock into one column or the other | |
So the committed answer sounds like: "This lock only dedupes thumbnail generation — worst case we render one twice, so a single Redis lock with a TTL is right, and Redlock would be complexity for nothing. But the settlement writer is a correctness lock — two holders means double-paying merchants — so that one lives in etcd, and the ledger table checks the fencing token on every write." One test, two locks, two different answers, each justified by its failure cost. That's the whole skill.
One more senior distinction, and it's the one that stops you from misusing everything above. Every system splits into a control plane — the part that decides who leads, where shards live, which config is active; it changes rarely — and a data plane — the part that serves every single user request. Consensus belongs in the control plane, and the reason is arithmetic, not taste.
Every consensus write is a network round trip to a majority plus a disk flush on each: single-digit milliseconds in one datacenter, and an intercontinental round trip — 80 to 150 milliseconds — if your quorum spans regions. Every write also serializes through one leader, which caps a healthy etcd cluster at a few tens of thousands of small writes per second, total, no sharding. Wonderful for elections and config. Hopeless in front of 500,000 requests per second — put etcd in your per-request path and you've installed a latency floor and a throughput ceiling in one move.
The correct wiring: the data plane watches the control plane and caches. Your app servers hold the shard map, the feature flags, and the current leader's address in local memory, refreshed by etcd watch notifications within milliseconds of a change. Requests never block on etcd. And when etcd goes down? The data plane keeps serving from its last-known-good cache. You temporarily lose the ability to change things — fail over a leader, ship a config — but not the ability to serve. Say that failure mode unprompted and you sound like someone who has operated this.
(The nuance worth one sentence at the whiteboard: databases like CockroachDB and TiKV deliberately put Raft in the write path — one small Raft group per shard, replicas placed close — accepting the majority round trip as the honest price of serializable SQL. Renting that trade from a database company is fine; rebuilding it by accident with etcd in your request loop is not.)
Pinterest wired ZooKeeper deep into its data plane: service discovery and application config flowed through it, with thousands of application servers holding live ZooKeeper connections. On paper that's safe — ZooKeeper is itself a replicated quorum system, not one machine. In practice, Pinterest's engineers wrote publicly that ZooKeeper had become a single point of failure for the site: a distributed service, but still one dependency, with failure modes (overload, session storms, its own outages) that were hard to prepare for — and when it had a bad day, the outage propagated to the whole site, because every service needed ZooKeeper answering to keep working. Their fix was not a better ZooKeeper. They decoupled: a small daemon on every host mirrors the needed ZooKeeper data to local disk, and applications read only the local copy. If ZooKeeper disappears, everything keeps running on last-known-good data; only the ability to push changes pauses. They later open-sourced the approach (as part of KingPin), running it across roughly 20,000 hosts with config updates delivered in seconds. The lesson is exactly this section: the coordination service belongs in the control plane, and the data plane must be able to survive its absence.
This chapter can feel bottomless — Raft alone has a PhD dissertation behind it — so here's the calibration. The L5 bar is fluent use of consensus as a primitive, plus its famous failure mode. It is not algorithm internals.
| You must volunteer (L5) | Garnish (nice, but staff-level or optional) |
|---|---|
| "I'd use etcd/ZooKeeper here" with the guarantee named: linearizable compare-and-swap, leases, watches | Raft internals: log matching, snapshotting, joint-consensus membership changes |
| Quorum math on demand: 2f+1 tolerates f; 3→1, 5→2; even sizes waste a node | Paxos variants (Multi-Paxos, EPaxos), flexible quorums |
| Lease-based election mechanics: TTL, heartbeat renewal, ephemeral nodes | Byzantine fault tolerance — nodes that lie, blockchain territory; irrelevant unless asked |
| The GC-pause/fencing failure, unprompted, the moment you propose any distributed lock | Deriving why randomized election timeouts prevent split votes forever |
| Consensus off the request path, with the numbers: majority round trip per write, tens of thousands of writes/sec ceiling | Designing a lease protocol's exact clock-skew margins |
| Efficiency vs correctness locks whenever Redis locks come up | Reciting the Redlock debate blow by blow |
And three answers that actively read as red flags: "I'll implement Raft in the service" (buys risk to reinvent a rentable primitive); "the lock service guarantees only one holder, so we're safe" (ignores pauses — no lock service can promise that alone); and "ZooKeeper means split brain can't happen" (ZooKeeper prevents split brain in the election; your stale ex-leader can still write unless something fences it).
Consensus rarely arrives as its own question at L5. It ambushes you from inside another design: a job scheduler, a failover story, a "how does the cluster know who's primary?" follow-up. A passing answer rents the primitive, states its guarantee, and raises the fencing failure before the interviewer does. Expect these probes, nearly verbatim:
Explain to a junior engineer, in five or six sentences, why a distributed lock with a timeout isn't enough to guarantee only one writer — and what a fencing token adds.
The lock has a timeout because a holder might die, and we can't tell a dead process from a slow one. But that timeout creates a trap: a process can freeze — say a long garbage-collection pause — hold the lock past its expiry, and then wake up with no idea any time has passed, so its very next instruction is a write it still believes is protected. Meanwhile the lock has legitimately moved to someone else, so two processes write at once. No check inside the frozen process can catch this, because the freeze can land between the check and the write. A fencing token fixes it by moving enforcement to the storage side: every lock grant comes with a number that only goes up, every write carries its number, and storage rejects any write with a number lower than the newest it has seen. The stale writer's own token betrays it — just like an old hotel keycard that stops working the moment the front desk issues a new one.
You're placing a 5-node etcd cluster across datacenters. Option A: 2 nodes in DC-East, 2 in DC-West, 1 in DC-Central. Option B: 3 in DC-East, 2 in DC-West. For each option, which single-datacenter failure stops the cluster from accepting writes? Which placement do you commit to?
Quorum is 3 of 5. Option A: lose East (2 nodes) → 3 survive, writes continue; lose West (2) → 3 survive, fine; lose Central (1) → 4 survive, fine. No single DC failure halts you. Option B: lose West → 3 survive, fine; but lose East and only 2 survive — no quorum, writes halt. Commit to A: it's the only placement where no single datacenter is load-bearing. The general rule: spread the cluster so that no failure domain contains a quorum-breaking number of nodes — which is why three failure domains is the magic minimum, and why two-datacenter consensus deployments are always quietly betting on one of the two.
Sketch leader election for a cron service ("run the nightly settlement job on exactly one of three schedulers") using etcd, including how you prevent a stalled ex-leader from double-running the job. Name the specific etcd features you'd use.
Each scheduler tries to create /election/settlement with a compare-and-swap ("create if absent"), attached to a ~10s lease it renews with keep-alives; losers watch the key and retry when it vanishes. The winner reads the lock key's mod_revision — a monotonically increasing number — and uses it as its fencing token. Every write the job makes to the settlement table carries that token, and the table enforces it: a token column plus a conditional update that rejects any write bearing a lower token than the highest recorded. Now the killer scenario is handled: leader 1 stalls 40 seconds, its lease expires, leader 2 elects with a higher revision and starts writing; when leader 1 wakes and writes, its stale token is rejected by the database, not by any code inside the frozen process. Bonus senior point: settlement writes should also be idempotent (p2c02), so even a rejected-then-retried run can't double-pay.
A teammate proposes reading feature flags from etcd on every request "so flags update instantly." The service does 30,000 requests/sec. What breaks, roughly when, and what do you build instead?
Two things break. Throughput: 30K reads/sec lands near a small etcd cluster's comfortable ceiling all by itself — before any other consumer — because linearizable reads involve the leader, and everything funnels through one Raft group. Latency: you've added a network hop with etcd's tail latency to every single request, and when etcd degrades (compaction, leader election, a slow disk), your entire service degrades with it — you've wired the control plane into the data plane. Build watch-and-cache instead: each app server loads flags at boot, holds them in memory (nanosecond reads), and subscribes to an etcd watch that pushes changes within milliseconds — "instant updates" was never actually the trade-off. If etcd goes down, servers keep running on last-known-good flags; you lose the ability to flip flags, not the ability to serve. That failure mode, stated unprompted, is the senior finish.
Classify each lock as efficiency or correctness, then commit to a tool for each: (a) "only one worker should send the weekly digest email batch", (b) "only one process may apply ledger postings for account 4471", (c) "only one node should refresh the shared cache entry when it expires."
(a) Efficiency. Two holders means some users get a duplicate email — embarrassing, not corrupting. A single Redis SET NX PX with a sane TTL is right; anything more is ceremony. (Even better: make sends idempotent per user-week, and the lock becomes pure optimization.) (b) Correctness. Two holders can double-post money — this is the fencing chapter in miniature: etcd/ZooKeeper election, fencing token from mod_revision or a sequential node, ledger table rejects stale tokens. (c) Efficiency — it's a cache stampede guard (p1c06); worst case two nodes recompute one value. Redis lock or even a local per-process guard suffices. The senior signal isn't the classifications alone — it's applying one test ("what happens with two holders?") three times and accepting three different amounts of machinery.
Priya and Arjun are editing the same design doc. The sentence on screen says the cat. At the same instant — same second, different laptops — Priya types fat before "cat", and Arjun deletes the word the at the front. Two tiny, innocent edits.
Each laptop applies its own edit instantly (nobody will use an editor where your own keystrokes lag), then sends the edit to the other side. Priya's edit travels as an instruction: "insert fat at position 4." Arjun's travels as: "delete 4 characters starting at position 0." Now watch what happens when each laptop applies the other's instruction, exactly as written.
Priya's laptop has the fat cat. It applies Arjun's delete — remove 4 characters from position 0 — and gets fat cat. Lovely. Arjun's laptop has cat. It applies Priya's insert — put fat at position 4 — but cat is only 3 characters long. Position 4 doesn't exist. Depending on the code, you get a crash, or the garbage string catfat . Two people made two reasonable edits, and now the "shared" document exists in two different versions, forever.
This is the concurrent-editing problem, and it is the crux of Google Docs, Figma, Notion, multiplayer whiteboards, and every offline-sync app you've ever used. There are exactly two famous families of solutions — OT and CRDTs — and at the senior level you're expected to know both, pick one for the scenario in front of you, and defend the pick. This chapter gets you there.
Before the two real solutions, kill the three "easy" ones, because interviewers love watching you kill them.
Lock the document. Old SharePoint and shared network drives did this: "This file is checked out by Bob." One writer at a time, zero merge problems, miserable product — collaboration was the feature, and a lock deletes the feature. Per-paragraph locks soften it, but two cursors in one sentence stays impossible.
Last write wins the whole document. Each save uploads the entire doc; latest upload replaces everything. Priya saves at 10:00:01, Arjun at 10:00:02 — every keystroke Priya made is silently gone. This is how naive file-sync tools eat homework. "Silently" is the killer word: nobody gets an error, the work just vanishes.
Timestamp every keystroke and apply them all in time order. Sounds rigorous. Two problems. First, clocks on different machines disagree (you saw clock skew in the replication chapter) — ordering by lying clocks is ordering by dice. Second, and deeper: even a perfect global order doesn't save you, because each operation's position was computed against the document as its author saw it. Arjun's "delete at position 0" was written against the cat. Replaying it against some other intermediate state changes what it deletes. Order was never the real problem — meaning is. An op's indexes are only valid in the context where it was born.
So the real requirement has two halves. Convergence: after everyone's edits reach everyone, all replicas show the same document. Intention preservation: that document contains what each person meant to do — Priya's word inserted before "cat", Arjun's "the " gone. Both solution families are answers to this pair. They just put the cleverness in different places.
Concurrent editing is hard because operations reference positions, and positions are relative to a document state that other people are changing at the same time. Every solution is a strategy for making an operation meaningful outside the context it was born in: OT rewrites the operation to fit the new context; CRDTs redesign the data so operations never needed rewriting in the first place.
Operational Transformation keeps the simple ops — insert-at-index, delete-at-index — and adds one crucial move: before applying a remote operation, transform it against the concurrent operations you've already applied, adjusting its positions so it still does what its author meant.
Run the opening disaster again, with transformation. Arjun's laptop holds cat (his delete already applied). Priya's insert(4, "fat ") arrives. The transform function reasons: "Arjun deleted 4 characters, all of them before position 4. So Priya's intended spot has slid left by 4." It rewrites the op to insert(0, "fat "), applies it to cat, and gets fat cat. Priya's side needed no rewrite: Arjun's delete covers positions 0 through 3, entirely before her insertion point at 4, so it applies as-is. Both replicas: fat cat. Converged, intentions intact.
OT is two proofreaders mailing correction slips for the same manuscript. A slip says "insert this word on line 12." If, before your slip arrives, the other proofreader deleted two lines above line 12, the editor's assistant adjusts your slip to "line 10" before applying it — because your intent was "before that particular sentence", not "line 12" as a magic number. The assistant who adjusts slips is the transform function. The analogy is honest about the hard part too: with many slips crossing in the mail, the assistant must adjust slips against slips against slips, and one wrong adjustment corrupts the manuscript for everyone.
With two users, transformation is a tidy square. With N users all editing concurrently, every op might need transforming against every other concurrent op, in a web of combinations that must all agree. The classic escape — invented at Xerox PARC in the Jupiter system (1995), inherited by Google Wave, and living on in Google Docs — is a central server that puts all ops into one single sequence. Each client then only ever diverges from the server by its own unacknowledged keystrokes, so it only ever transforms along one axis: my pending ops versus the server's stream. The nightmarish N-way case never happens.
That design choice explains OT's whole personality:
Google Wave (2009) was Google's bet that email, chat, and docs should be one realtime product, built on OT descended from the Jupiter paper. Joseph Gentle, an engineer who worked on Wave in Sydney, wrote the sentence OT is now famous for: "Unfortunately, implementing OT sucks. There's a million algorithms with different tradeoffs, mostly trapped in academic papers. The algorithms are really hard and time consuming to implement correctly… Wave took 2 years to write and if we rewrote it today, it would take almost as long to write a second time." Wave died in 2010 — a product failure more than an algorithm failure — but the lesson stuck: the transform machinery consumed elite engineers for years. Gentle spent years rebuilding the tech as the open-source library ShareJS so nobody would have to write it again — then in 2020 published a widely read essay titled "I was wrong. CRDTs are the future," arguing a decade of CRDT research had closed the performance gap. When one of OT's best implementers migrates camps, that's data.
So OT works, ships, and scales — Google Docs is the existence proof. But its two structural needs, a central server and mostly-connected clients, are exactly the assumptions the modern world keeps breaking: offline-first mobile apps, peer-to-peer tools, local-first software that treats the server as optional. For those, you want merges that need no referee at all.
A CRDT — Conflict-free Replicated Data Type — is a data structure built so that merging any two replicas always produces the same result, no matter what order the merges happen in, no matter how many times you repeat them. No transform functions, no referee, no coordination. Every replica edits locally, replicas swap states or ops whenever they happen to meet, and math — not a server — guarantees they converge.
The trick is choosing merge rules with three properties you already know from arithmetic: the merge must be commutative (A merged with B = B merged with A), associative (grouping doesn't matter), and idempotent (merging the same thing twice changes nothing). max() has all three: max(3, 7) = max(7, 3), and max(7, 7) = 7. Build your structure out of operations like that, and replicas can gossip states in any order, with duplicates and delays, and still agree. That last property, idempotence, is quietly the most practical one: on a flaky network you will deliver the same update twice, and a CRDT simply doesn't care.
G-counter (grow-only counter). You'd think a distributed counter is trivial — it isn't, because "add 1" messages get lost and duplicated. The CRDT answer: each replica keeps its own slot in a map — {a: 3, b: 1} means replica-a incremented 3 times, replica-b once. You only ever increment your own slot. Merge = element-wise max. Total = sum of slots. Because each slot only grows and only one writer touches it, max is always safe. A PN-counter is two G-counters (one for increments, one for decrements); value = P minus N. This is how Riak shipped mergeable counters, and it's the right shape for likes, view counts, and metrics across regions.
Three teachers count students on a field trip. Bad plan: shout running totals at each other and try to add them up — totals crossing mid-air get double-counted. CRDT plan: each teacher counts only her own group on her own tally sheet. Whenever any two teachers meet, they photocopy sheets and each keeps, per teacher, the highest tally seen. Meet in any order, photocopy twice by accident — harmless. The grand total is always the sum of the per-teacher tallies, and everyone converges on it without a head teacher coordinating anything. The "one writer per slot" rule is what makes max safe — that's the part the analogy must keep.
LWW-register (last-writer-wins register). A single value plus a timestamp; merge keeps the higher timestamp. It's the simplest possible CRDT and the most dangerous one, because its conflict "resolution" is silent data loss by design: when two people write concurrently, one write evaporates, no error, no trace. Add clock skew and it gets worse — a device with a fast clock stamps its writes "in the future" and beats edits that genuinely came later. Real sync systems have shipped this bug: a phone with a wrong clock whose stale settings kept overwriting fresh ones. LWW is a fine choice when losing one of two concurrent writes is acceptable (a mouse cursor position, a "who's online" flag, one property of one design object). It is an outrageous choice for anything a human typed and expects to keep.
OR-set (observed-remove set). Sets need a policy for the classic race: Priya removes "milk" from the shared list while Arjun, concurrently, re-adds it. The OR-set decides adds win, via a neat mechanism: every add gets an invisible unique tag, and a remove deletes only the tags it has observed. Arjun's concurrent add carries a fresh tag Priya's remove never saw, so it survives the merge. You could build a remove-wins set instead — the point a senior makes out loud is that "conflict-free" doesn't mean conflicts vanish; it means you chose the winner in the data structure's design instead of at merge time. Amazon's original Dynamo shopping cart got this wrong in the other direction, famously resurrecting deleted items when merging carts — the anomaly OR-sets were designed to fix.
Sequence CRDTs — text at last. Text is the boss level, because text is a sequence and sequences are all about positions — the very thing that broke in the cold open. The family of solutions (RGA, the "replicated growable array," is the classic algorithm; Yjs and Automerge are the production libraries) shares one core idea: kill the index. Every character gets a permanent unique ID at birth, and an insert says "I go after character #a17", not "I go at position 4". IDs never shift when other text changes, so concurrent inserts don't tread on each other; tie-break rules order simultaneous inserts at the same spot consistently everywhere. You don't need the internals — you need the idea (position by identity, not by index) and the two costs it buys them:
| Dimension | OT | CRDT |
|---|---|---|
| Central server | Required — it defines the op order | Optional — any topology, P2P included |
| Offline editing | Minutes are fine; long divergence is painful | Native — merge whenever, however late |
| Realtime text | Battle-tested for 15+ years (Docs) | Practical now via Yjs/Automerge; costs metadata |
| Doc/memory overhead | The doc stays just the text; the cost is a buffer of unacknowledged ops, bounded by round-trip time rather than by edit history | IDs + tombstones; small multiple of text with modern libs |
| Correctness burden | In the transform functions (hard to write, silent when wrong) | In the data-type design (hard to invent, but you use a library) |
| Non-text state (counters, sets, presence) | Awkward fit | Natural fit — this is home turf |
| Global invariants ("seat sold once", "balance ≥ 0") | Neither. Invariants need coordination — locks, transactions, consensus. Merging freely and enforcing a global rule are opposites. | |
The last row is a favorite senior trap. CRDTs guarantee convergence, not correctness of business rules. Two replicas can each sell the last concert seat while partitioned, and the merge will converge beautifully on a state where the seat sold twice. If the requirement is an invariant that must hold globally, the answer is coordination (Chapter 2.3), not a cleverer merge.
The honest decision rule, committed and justified: realtime collaborative text with a server you control → OT or a server-based sequence CRDT, both defensible; say why you picked yours. Offline-first, peer-to-peer, or multi-leader sync → CRDT, no contest. Non-text shared state — counters, sets, presence, settings → small CRDTs, they're practically free. Global invariants → neither; coordinate. And there's a third path the next story earns: if you have a central server and your data isn't text, you may not need either algorithm at full strength.
When Figma built multiplayer editing (shipped 2016; detailed in Evan Wallace's 2019 post "How Figma's multiplayer technology works"), they studied OT first — it's what Google Docs uses — and rejected it as more complexity than a design tool needed. But they didn't adopt full CRDTs either. A Figma document is a tree of objects, each a bag of properties, and their protocol is CRDT-inspired: concurrent changes to different properties of one object both survive; concurrent changes to the same property resolve last-writer-wins, with the central server — which every edit flows through — deciding "last" by arrival order, so no clock-skew roulette. That's the LWW register's data loss, accepted with eyes open: if two designers set the same rectangle's fill in the same instant, one fill wins, and the loser sees it happen and can re-apply in a click. The same trade on prose would eat sentences. Because the server is authoritative, they skipped the machinery decentralized CRDTs need; child ordering uses fractional indexing — positions as fractions between neighbors — rather than sequence-CRDT lists. The lesson interviewers want you to extract: Figma scoped the problem — design objects, central server, property granularity is enough — then bought exactly as much merge machinery as that problem required, not a bit more.
Whichever family you pick, the surrounding system is the same, and mentioning it is cheap senior signal because it shows you've thought past the algorithm to the service:
Calibration matters here more than in most chapters, because this topic has a bottomless academic end and candidates regularly drown in it. Nobody — nobody — expects you to derive a CRDT or write a transform function at a whiteboard. Here's the actual L5 bar:
You must volunteer, unprompted:
the cat square from this chapter is enough. This is the single highest-value minute in the whole topic: it proves you understand the problem, not just the vocabulary.Staff-only garnish — name-drop honestly if it fits, never pretend depth: tombstone GC and the causal-stability condition that makes it safe; version vectors and causal ordering internals; OT's TP2 property and why P2P OT died of it; sequence-CRDT interleaving anomalies; byzantine (malicious-peer) settings. A perfect senior sentence: "Tombstone GC needs to know every replica has seen the delete — that's causal-ordering machinery I'd pull a library for, not build."
This topic appears two ways: headline ("Design Google Docs / a collaborative whiteboard") or ambush (you're designing something else, and the interviewer adds "now two users edit the same item offline"). A passing senior answer names both families in under a minute, commits with a scenario-tied reason, walks one merge example concretely, and volunteers where the design loses data or grows without bound. Probes you should expect verbatim:
Explain to a junior, in five sentences, why Google Docs doesn't corrupt the document when two people type in the same sentence — and why their approach struggles offline.
Every keystroke becomes a tiny operation like "insert 'x' at position 12," and the problem is that your position 12 was computed before my simultaneous edit moved everything. Google Docs solves it with operational transformation: a central server puts all operations in one agreed order, and each incoming operation is rewritten — its positions shifted — to account for the concurrent edits already applied, so everyone converges on the same text with both people's intent preserved. That rewriting only stays simple because the server provides one official sequence to transform against. If you go offline for a week, you come back with thousands of operations that must be transformed against thousands of everyone else's, and the scheme gets fragile — OT is built for editing together, not apart. That's why offline-first tools use CRDTs instead, where every character carries a permanent ID so positions never shift and replicas can merge in any order without a referee.
Transform drill. Document: hello. Concurrently, A does insert(5, "!") and B does delete(0, 1) (removing the h). Show each replica's final string with correct transformation, and name which op needed rewriting and why.
Intent: A wants ! after the last letter; B wants the h gone. Converged answer: ello!. A's replica: hello!, then B's delete applies unchanged (its range [0,1) is entirely before the insert) → ello!. B's replica: ello, then A's insert must be rewritten: B deleted 1 character before position 5, so shift left by 1 → insert(4, "!") → ello!. Only A's op needed transforming, because only its position sat after the region B changed. If you applied A's op raw at position 5 on the 4-character ello, you'd be writing past the end — the cold-open bug again.
Commit-and-justify: (a) a note-taking app for field geologists who sync laptop-to-laptop by local Wi-Fi after two weeks off-grid; (b) a browser-based pair-programming editor your company hosts; (c) a global "total downloads" counter shown on your homepage, incremented from five regions. Pick OT or CRDT (or neither) for each, one sentence of justification.
(a) Sequence CRDT (Yjs/Automerge): peer-to-peer merge after long divergence is CRDT home turf and OT's worst case — there may not even be a server. (b) Either is defensible with a server you control; a committed answer: "OT-style with server ordering — clients are online while editing, ops stay tiny, and the server I already run gives me the total order for free; I'd reach for Yjs instead if offline editing enters the roadmap." (c) PN-counter (or G-counter if it never decrements): counters aren't text, need no server coordination, and per-replica slots merged by max survive partitions and duplicate delivery. Bonus senior point on (c): if the product ever needs "exactly N," a counter CRDT gives convergence, not an invariant.
What breaks: a mobile app syncs user settings as LWW-registers keyed by device wall-clock time. Describe the specific silent failure when a user's phone clock is 30 minutes fast, and two mitigations of different weights.
The fast phone stamps every write 30 minutes in the future. The user changes a setting on the phone, then changes it again on their laptop ten minutes later — the laptop's genuinely newer write carries a smaller timestamp and loses every merge. Worse, it's sticky: until real time passes the phone's stamp, no other device can update that setting; edits keep silently reverting, and support can't reproduce it. Lightweight mitigation: stamp with a server-assigned time or hybrid logical clock on sync instead of raw device clocks (Figma's version: the server's arrival order defines "last"). Heavyweight: stop using LWW for user-visible data — keep per-field versions or a small CRDT and surface true conflicts. Detection tip worth saying aloud: log every LWW discard with both timestamps; a discard whose "loser" is newer in server time is your smoking gun.
G-counter drill. Replicas: A = {a:5, b:2, c:0}, B = {a:4, b:6, c:1}, C = {a:5, b:6, c:3}. What's the merged value? Why is per-slot max safe here, and what single rule violation would make the counter wrong?
Element-wise max = {a:5, b:6, c:3} → value 14. Max is safe because each slot has exactly one writer and only ever grows, so the largest value seen per slot is that replica's true count — duplicates and reordering can't inflate it (idempotent), and no increment can be lost (a bigger value always survives). The rule you must not break: incrementing someone else's slot. If replica A ever bumps slot b, two writers race on one slot and max starts silently discarding real increments. Same reason you can't merge by summing slots: merging twice would double-count.
Sketch the reconnect path for a collaborative whiteboard: a client returns after 3 hours offline with 200 local ops; the document advanced 5,000 server ops meanwhile. Name the pieces and the order of events, including how you avoid applying anything twice.
Client state: last-seen server position (say op 41,000) + 200 pending ops, each carrying a stable ID (clientID, seq). Reconnect: (1) client sends its position; (2) server streams the missing tail — or, if the tail is huge, the latest snapshot plus ops since it; (3) client integrates the 5,000 remote ops — CRDT: just merge, IDs make positions unambiguous; OT: transform its 200 pending ops over the incoming stream as it applies them; (4) client submits its 200 ops; server appends any ID it hasn't logged and drops the rest — that dedupe is what makes a mid-upload disconnect and retry harmless; (5) server broadcasts the accepted ops; the origin client recognizes its own IDs and skips reapplying. Volunteer the failure mode: after a server deploy, thousands of clients hit step 2 at once — a reconnect storm — so tail-serving must be cheap (snapshots) and jittered.
It's 6:47 on a Tuesday morning and your phone will not stop. Every alert you own fires at once — API error rate, database timeouts, dead queue workers, all screaming together. You open your cloud provider's status page and there it is, in careful corporate prose: "We are investigating increased error rates in us-east-1." Not a server. Not a rack. Not even an availability zone. The region — the entire geographic cluster of data centers where your company's software lives — is having a very bad morning.
So here is the question this chapter exists to make you answer: what does your app do right now? Not what the architecture diagram implies. What actually happens. For most companies the honest answer is: nothing. It's down, and everyone waits for someone else's engineers to fix it. All those replicas from Part 1, the load balancers, the carefully sharded database — every copy lives in the same region. You built beautiful redundancy, then put all of it in one basket the size of Northern Virginia.
Region-level failure is rare, but it is not hypothetical. In February 2017, a mistyped command during routine S3 debugging took S3 in us-east-1 down for about four hours — and a startling fraction of the internet with it. In December 2021, a us-east-1 network incident degraded services from Slack to Disney+ for the better part of a day. If your design review has never asked "what happens when the region dies?", the answer is being decided for you, and the answer is "we go down with it."
This chapter covers the two big ideas that answer that question: multi-region architecture — copies of your system in different geographies — and cell-based architecture — walls inside each geography, so one failure can't take the whole fleet. They're the same idea at two different scales: control the blast radius.
Multi-region is expensive and complicated, so be honest about why you'd do it. There are exactly three reasons, and each pushes you toward a different design.
1. Physics. Light in fiber travels at roughly 200,000 km per second — two-thirds of its speed in vacuum. That sounds fast until you do the math on a round trip. A request from Sydney to a server in Virginia covers about 16,000 km each way: 160 milliseconds of pure travel time, before a single line of your code runs. Real networks add routing overhead on top.
| Path | Typical round trip |
|---|---|
| Same city / same region | 1–2 ms |
| Virginia ↔ Oregon (cross-US) | ~70 ms |
| Virginia ↔ London | ~80 ms |
| Virginia ↔ Mumbai | ~190 ms |
| Virginia ↔ Sydney | ~200 ms |
And it compounds: setting up a fresh HTTPS connection takes roughly three round trips before the first real byte moves. For your Sydney user, that's ~600 ms of nothing. No code optimization fixes this — only moving servers closer to the user does. A CDN (Part 1) moves your static content closer. Multi-region moves your servers and data closer.
2. Availability. An availability zone is a separate building with separate power; a region is a cluster of them. Multi-AZ protects you from a building fire. But a region still shares things: the provider's regional network, its control planes, its deploy pipeline. When those break, every AZ breaks together. The only defense against "the provider is having a bad day here" is having a there.
3. Law. Data residency rules say certain data about certain people must be stored — sometimes processed — inside a specific country or bloc. European privacy law, India's rules for payment data, and many national variants all push this way. Residency forces region placement no matter what latency and availability say, and it constrains replication: you can't quietly copy the EU users' database to Virginia "for safekeeping" if the law says it stays in the EU.
The cheapest version of multi-region is the one where only one region does the real work. Start there.
In an active-passive setup, one region (the active, or primary) serves all traffic. A second region (the passive, or standby) continuously receives a copy of the data — the async replication machinery from Part 1, stretched across a continent — and otherwise sits quiet. When the primary dies, you promote the standby and point traffic at it.
Standbys come in degrees of readiness, and each degree costs more: backup-and-restore (just backups in another region; hours to rebuild), pilot light (data replicating live, but only a skeleton crew of servers running), warm standby (a full but scaled-down copy of the stack), and hot standby (full capacity, ready in minutes). You are buying down recovery time with money.
Two acronyms carry this entire topic, and interviewers expect you to use them with numbers attached. RPO — Recovery Point Objective — is how much recently written data you accept losing, measured in time. RTO — Recovery Time Objective — is how long until you're serving traffic again. Concretely: your cross-region replication runs asynchronously with 2–5 seconds of lag, because making every write wait ~70+ ms for a cross-country confirmation would be brutal (that's the CAP trade-off from Part 1 wearing a business suit). So when the region dies, the last few seconds of writes — writes your users already saw succeed — existed only in the dead region. They are gone. Async replication means a nonzero RPO. That is not a bug in your setup; it is the deal you signed.
The senior question is never "what's our RPO?" — it's "which writes may we lose?" One number for the whole system is a junior smell, because the honest answer differs by data class. Losing five seconds of session tokens: nobody notices, they log in again. Five seconds of analytics events: shrug. Five seconds of shopping-cart edits: mildly annoying. Five seconds of confirmed payments: you are writing apology emails and reconciling money by hand. The E5 sentence is: "Everything rides async replication with an RPO of a few seconds, except the payments ledger, which is synchronously replicated — those writes pay the cross-region latency, and that's the right trade." Different data, different promises, stated unprompted.
A shopkeeper photocopies her ledger and sends the copies by courier to her sister's shop across town every few minutes. If her shop burns down, the pages written since the last courier run are ash — that's the RPO. The time it takes to unlock the sister shop, brief the staff, and open the doors — that's the RTO. Want zero pages lost? Then the courier must carry every page the moment it's written, and she can't serve the next customer until he's back — that's synchronous replication, and it makes every sale slower. Where the analogy breaks: real replication streams continuously, so the "courier gap" is seconds of lag, not minutes.
Walk the numbers like an owner: replication lag ~3 s. Detection to confident alarm, 3 minutes — you must distinguish a blip from a death, and paging a human on every blip destroys trust. A human confirms and pulls the trigger at minute 6, because fully automatic region failover that fires on a false alarm creates worse problems (coming shortly). Traffic shift plus database promotion, 6 more minutes. Caches warm by minute 15. Result: RTO ≈ 15 minutes, RPO ≈ 3 seconds. Now the numbers drive decisions: if the business can't eat 15 minutes, you pay for a hot standby and rehearsed automation. If it can't eat 3 seconds of lost orders, you synchronously replicate the order table — just that table — and every checkout gets ~80 ms slower. Say the cost out loud when you commit.
But notice what active-passive doesn't fix: your Sydney users still pay 200 ms on every request, and an entire region's worth of hardware sits idle. Which raises an obvious, dangerous idea — why not make both regions work?
In an active-active setup, every region serves traffic. Users hit the nearest one, writes happen everywhere, and regions replicate to each other in both directions. Latency: solved. Idle hardware: none. Failover: barely an event — the survivors just absorb the traffic. It sounds strictly better than active-passive, and it isn't, because of one word: conflicts.
Both regions did their job. Both writes were acknowledged. Now the replication streams cross, and each region holds a different truth. You have three families of answers, in ascending order of effort:
There's a fourth answer, and it's the one that wins most real designs: refuse to have the fight at all. Assign every user a home region — usually the one nearest where they signed up. All writes for that user route to their home region, no matter where they are today. Reads are served everywhere, from local async replicas. One writer per user means write conflicts are structurally impossible — you didn't resolve the conflict, you made it unrepresentable.
The costs are honest and small: travelers pay write latency to their far-away home. Read-your-own-writes needs care — right after Priya saves, the EU replica may not have her edit yet, so you briefly pin her reads to her home region or to her session (the same replication-lag trick from Part 1). And a home region dying still needs the full active-passive failover story — for that slice of users. Early Facebook ran the global version of this for years: every write in the world flew to California's master databases; reads came from local replicas. It's not glamorous. It works.
So commit with a ranking, not a shrug: active-passive when your users live in one geography; read-local-write-home when reads are global and writes are per-user; true multi-master active-active only for the narrow slices of data with a well-defined merge — counters, presence, carts. Naming which data gets which treatment is precisely the senior move.
Two mechanisms route users to regions, and they fail differently.
Geo-DNS: the DNS server gives different answers depending on where the question comes from — Sydney resolvers get the Sydney region's IP. Add health checks and it doubles as your failover switch: mark a region unhealthy, and DNS stops handing out its address. The weakness is that DNS answers are cached, and the TTL — the "expire after" sticker on each answer — is advisory. Resolvers, corporate proxies, and long-running app processes routinely hold IPs long past TTL. After a failover, a stubborn tail of traffic keeps arriving at the corpse for minutes to hours. Plan for the tail: clients should retry against an alternate endpoint, not just trust DNS.
Anycast: the same IP address is announced from many locations at once, and the internet's routing system (BGP) delivers each packet to the nearest one. There's nothing to cache stale — withdraw the announcement in the dead region and traffic reroutes in seconds. This is how big CDNs and DNS providers front their networks. The pragmatic stack: anycast at the edge for instant shifts, geo-DNS above it, health checks driving both.
At the whiteboard, "my region died" should trigger a rehearsed six-step script, delivered unprompted. Here it is; learn it as a sequence, because the order matters.
Ask any team with a disaster-recovery plan one question: when did you last actually fail over? If the answer is "never," they don't have a DR plan — they have a DR document. Real failovers surface the things no document predicts: the standby is sized for 10% of production traffic; an IAM permission exists only in the primary; a hostname is hardcoded; a critical cron job runs nowhere else. And the cruelest one — during a genuine regional outage, every other customer of your cloud provider is evacuating into the same neighboring region at the same time, and the spare capacity you assumed isn't there.
The only fix is rehearsal. GameDays — scheduled drills where you deliberately break production-like systems, up to and including evacuating a region — turn failover from a heroic first-time event into a boring routine. If you're not willing to run the drill, you already know the failover doesn't work.
On Christmas Eve 2012, an AWS maintenance process accidentally deleted production state data for Elastic Load Balancers in us-east-1. Netflix ran essentially everything in that one region — and streaming died on TVs across the Americas on one of the biggest movie nights of the year, for hours, over something Netflix's own engineers could not touch. The response reshaped the company: Netflix rebuilt its user-facing path to run active-active across multiple AWS regions, so any region can absorb another's users. Then the properly senior part: they built Chaos Kong, a tool that deliberately evacuates an entire region in production — on a regular schedule, not during emergencies — to prove the failover works. Because they'd rehearsed region evacuation many times before ever needing it in anger, shifting a real failing region's traffic became a practiced routine instead of a bet. That's the whole GameDay philosophy in one company: the outage you've rehearsed is an inconvenience; the one you haven't is an incident.
Now for an uncomfortable observation. Multi-region protects you from Virginia. It does not protect you from yourself. Most outages aren't geography — they're a bad deploy, a poison request that crashes whatever server touches it, one whale customer whose traffic tips over shared infrastructure. Multi-region does nothing for these, because your replication and your deploy pipeline lovingly deliver the same poison to every region you own.
Enter cell-based architecture. A cell is a self-contained mini-stack — its own load balancer, its own app servers, its own database — serving a fixed subset of customers. Cells share almost nothing; a thin routing layer in front maps each customer to their cell and stays out of the way. Now a bad deploy or a poison customer takes down one cell: 5% of customers, not 100%. The fleet's blast radius shrinks to the size of the wall you built.
Cells buy you two quieter superpowers too. Deploys roll cell by cell, so the first cell is a natural canary with a real, bounded user population. And cells have a fixed maximum size — you scale by stamping out more cells, not growing them — which means no cell ever operates at a scale you haven't already tested. Growth stops being a journey into the unknown.
Ships stopped sinking from single leaks when builders added bulkheads — watertight walls splitting the hull into compartments. Breach one compartment and it floods; the ship floats on. Cells are bulkheads for your fleet. And the analogy carries its own warning, courtesy of the Titanic: her bulkheads didn't reach full height, so water spilled over the top from one compartment into the next. Your version of water-over-the-top is any shared component spanning all cells — the cell router, a shared control plane, a global database. Whatever crosses the wall can sink every compartment, so keep that layer tiny, dumb, and boring.
One problem remains. With plain cells, each customer lives in exactly one cell — so the poison customer's innocent cellmates are 100% down while the rest of the fleet hums. Can we protect even them?
Instead of assigning each customer one cell, deal each customer a small random hand of them — say 2 cells out of 8. Requests try either cell; clients retry against the other on failure. Now trace a poison customer: they hurt their two cells. But their cellmates each hold a different second cell, so every innocent customer retries onto healthy capacity. A customer is fully down only if their entire hand is down — and with 8 cells choose 2, there are 28 distinct hands, so barely anyone shares your exact fate. Scale the numbers up and "barely anyone" becomes "statistically no one."
DNS is the internet's most attractive DDoS target: knock out a provider's name servers and every site they serve goes dark. Amazon's Route 53 answered with shuffle sharding, and the numbers are worth memorizing because they make the idea vivid. Route 53's capacity is divided into 2,048 virtual name servers. Every customer domain is dealt a hand of four of them — and with 2,048 choose 4, there are about 730 billion possible hands. AWS additionally guarantees no two customer domains ever share more than two of their four servers. So when an attacker floods some domain, that customer's four name servers take the beating — while essentially every other customer has a mostly-different hand, and DNS resolvers automatically retry across a domain's servers by design. An attack sized to destroy "the fleet" ends up isolating roughly one customer: the target. This pattern, spelled out by Amazon in its Builders' Library, is a core reason Route 53's query-answering data plane carries a 100% availability SLA — and why AWS now bakes cells and shuffle sharding into service designs across the company.
On June 30, 2021, network hardware in one AWS availability zone started intermittently dropping traffic — a gray failure, sick rather than dead, the kind automated health checks are worst at catching. Slack's services were spread across AZs with cross-AZ calls everywhere, so one flaky zone degraded the whole product while dashboards mostly shrugged. The engineers' postmortem wish was disarmingly simple: a button that tells all systems "this AZ is bad; avoid it." So over roughly a year and a half, Slack rebuilt its critical user-facing services into a cellular architecture, described in detail on its engineering blog in 2023: each cell is aligned to one availability zone and keeps its traffic inside itself ("siloed"), so a sick zone's weirdness stops at the cell wall. The button exists now — it targets removing traffic from a sick zone within five minutes, and when the zone recovers, operators can send back as little as 1% of traffic to test it before trusting it again. Note what Slack's version teaches: cells don't have to be region-sized. The bulkhead pattern works at whatever scale your failures come in.
Multi-region is a topic where candidates routinely miscalibrate in both directions — hand-waving "we'd go multi-region" like a magic phrase, or diving into BGP minutiae nobody asked for. Here's the honest line between what an L5/E5 answer must volunteer and what is staff-level garnish.
| L5 must volunteer, unprompted | Staff-only garnish (fine to mention, never required) |
|---|---|
| The region-death script — detection, shift, promotion, fencing, reconciliation — before being asked | Cell sizing math and cell-migration mechanics |
| Concrete RPO/RTO numbers for this design, and which writes sit in the loss window | Cross-cell routing-layer and control-plane design |
| Per-data-class choices: "sessions async, ledger sync" — with the latency cost stated | Cross-region egress cost modeling in dollars |
| A committed topology for this system — active-passive, read-local-write-home, or active-active — with the reason | BGP/anycast internals beyond one sentence |
| The conflict story, if you proposed multi-writer anything | CRDT merge semantics in depth (previous chapter's territory) |
| "And we drill this" — one sentence on GameDays | Shuffle-sharding combinatorics beyond the idea |
And the inverse is also senior signal: knowing when not to. A seed-stage product with users in one country needs one region, multi-AZ, tested backups, and a written restore-time estimate — proposing active-active for it reads as buzzword architecture, exactly the failure mode from chapter one. Multi-region is bought with real money and real complexity; the senior move is naming the trigger that would make you buy it.
Multi-region rarely arrives as its own question. It arrives as a grenade tossed into your finished design: "Looks good — now us-east-1 just went dark." A passing senior answer commits to a topology for this specific system, attaches RPO/RTO numbers per data class, walks the death script calmly, and mentions drills in one breath. Probes you should expect verbatim:
Explain RPO and RTO to a junior engineer in 4–5 sentences, using an online store as the example — and make sure they understand why async replication means RPO can't be zero.
RTO is how long the store is down after a disaster: if our region dies at noon and we're selling again at 12:15 from the backup region, RTO was 15 minutes. RPO is how much recent data the disaster erases: the backup region receives its copy of every order a few seconds after the primary confirms it, so the orders from those last few seconds die with the primary — customers saw "order confirmed," and we have no record. That gap exists because our replication is asynchronous: the primary tells the customer "success" first and copies to the other region a moment later, which keeps checkout fast. Making RPO truly zero means the primary can't say "success" until the far region also has the order — every checkout then waits for a cross-country round trip. So we choose per data type: a few seconds of lost sessions is fine, but for payments we pay the slow synchronous path, because losing money isn't a latency optimization.
Your e-commerce stack: primary in Mumbai, warm standby in Singapore, async replication with p99 lag of 5 seconds, DNS TTL 60 s, and a manual failover runbook that takes about 20 minutes end to end. State the RPO and RTO. Then the business demands RTO under 5 minutes and zero lost payments. What exactly changes?
Today: RPO ≈ 5 seconds of acknowledged writes (the replication lag — those orders exist only in Mumbai when it dies), RTO ≈ 20+ minutes (runbook time, plus a stale-DNS tail beyond the 60 s TTL still hitting Mumbai). For RTO < 5 min: hot standby at full capacity, automated promotion behind well-tested detection, and traffic shifting that doesn't depend on DNS caches — anycast or client-side failover — plus regular drills, because untested automation is fiction with extra steps. For zero lost payments: synchronous replication for the payments ledger only — each payment write waits ~70–90 ms for Singapore's acknowledgment before confirming. Everything else stays async. Notice the shape of the answer: you didn't make the whole system RPO-zero; you named the one data class that needs it and paid the latency only there.
A food-delivery app runs read-local-write-home across three regions. Classify these data classes and commit to a replication treatment for each: (a) live order state ("driver is 2 minutes away"), (b) restaurant menus, (c) user sessions, (d) the payments ledger, (e) analytics click events.
(a) Live order state: an order happens in one city — pin it to that region entirely; cross-region replication adds lag to the most latency-sensitive data for no benefit. If the region dies mid-delivery, accept that in-flight orders degrade; that's an explicit, stated loss. (b) Menus: write-rarely, read-everywhere — async replicate globally; seconds of staleness on a menu is harmless. (c) Sessions: async, RPO of seconds is fine — worst case a user logs in again. (d) Payments ledger: synchronous replication to a second region, RPO zero, and single-writer (no multi-master money, ever). (e) Analytics: fire-and-forget through a queue; losing seconds of events during failover is explicitly acceptable. The senior part isn't the choices — it's that each class got its own sentence with its own accepted loss.
Your automated failover fires on a false alarm: the primary region was healthy, just briefly unreachable from the monitor. Walk through what goes wrong, then name the two guards that prevent it.
The standby promotes and starts taking writes. Meanwhile the "dead" primary is alive, still receiving the stale-DNS tail of traffic, still accepting writes. Two primaries, two diverging histories — split brain. Every hour it persists makes reconciliation harder, and some conflicting writes (two sales of the same last item) have no clean merge. Guard one: detection quorum — declare a region dead only when multiple independent vantage points agree, for a sustained window, on user-visible symptoms; a single monitor's opinion never pulls the trigger. Guard two: fencing — promotion bumps a configuration epoch, and every downstream system rejects writes stamped by the old leader, so even a live-but-deposed primary can't commit anything. Then reconciliation still happens for the seconds before fencing took effect — which is why GitHub's 2018 incident cost 43 seconds of partition and over 24 hours of careful cleanup.
You run 8 cells with customers shuffle-sharded onto hands of 2 cells each. A poison customer's requests crash any server that touches them. Compare the blast radius against plain one-cell-per-customer assignment. How many distinct hands exist, and what does a cellmate of the poison customer experience?
Plain cells: the poison customer takes down their cell; every cellmate — about 1/8 of all customers — is 100% down. Shuffle sharding: the poison customer crashes both cells in their hand, so 2 of 8 cells are down (a real cost — shuffle sharding trades slightly wider partial damage for far less total damage). There are 8-choose-2 = 28 distinct hands. A customer sharing one cell with the poison hand still has their other cell healthy: their retries land there, so they run degraded, not down. Only a customer dealt the exact same two cells is fully out — 1 in 28 odds per customer, and at Route 53's scale (2,048 choose 4 ≈ 730 billion hands, max 2 shared) the chance of full overlap rounds to zero. The pattern's whole value is in that shift: from "some customers are completely down" to "almost everyone is merely a retry away from fine."
You did everything right. Chapter 1.9 taught you consistent hashing, and you used it: 64 partitions, keys spread beautifully, each partition holding almost exactly 1/64th of your users. The dashboards have been green for six months. You could frame them.
Then a K-pop star joins your platform. On Tuesday evening she posts a photo, and forty million people open it within an hour. Your phone buzzes. p99 latency is up 40x. Timeouts are climbing. And here's the part that makes the on-call engineer question reality: the database fleet's average CPU reads 11%. The fleet is bored. Sixty-three partitions are napping — and partition 23, the one that happens to hold her data, is pinned at 100%, its request queue growing by the second.
Nothing is wrong with your hashing. Hashing distributes keys evenly. Nobody promised it would distribute traffic evenly — and today, the entire internet cares about exactly one key.
This is the hot partition problem, and its most famous costume is the celebrity problem. It is one of the most reliably asked senior topics there is, because it sits at the exact spot where textbook sharding meets the real world. The moment an interviewer says "one of your users has 100 million followers," they are ringing a bell, and you are expected to come running with a very specific set of answers. This chapter is that set.
Every sharding scheme you learned — consistent hashing from chapter 1.9 included — makes a quiet assumption: each key gets roughly the same amount of attention. Reality never cooperates. Real traffic follows a power law — a distribution where a tiny fraction of items gets a huge fraction of the attention. A few accounts have millions of followers while the median has a hundred. One post goes viral while a billion others get four likes. One enterprise tenant in your B2B product generates 40% of all API calls. One product page carries the flash sale.
The numbers are brutal. On social platforms it's routine for the top 1% of keys to receive well over 90% of reads. So "perfectly balanced by key count" and "perfectly balanced by load" are different universes. Your hash function solved the first. The second is this chapter.
A library shelves its books alphabetically — a beautifully even split, every shelf holding about the same number of books. Then a new Harry Potter comes out. Suddenly three hundred people crowd one shelf while the rest of the library sits empty. The shelving system isn't broken; it balanced books, not readers. No rearrangement of the other shelves helps, because the problem isn't where the books are — it's where the people are. That's a hot partition: the storage is balanced, the demand is not. (Where the analogy breaks: a library can put out fifty copies of one book in a day. Making "more copies" of a partition that's absorbing writes is the hard part — copies of data have to be kept in sync.)
And note the shape of the problem carefully, because the fix depends on it: a key can be read-hot (forty million people opening one profile), write-hot (eight thousand likes per second landing on one counter), or both. Keep asking "read-hot or write-hot?" throughout this chapter — the mitigation ladder forks on exactly that question.
Here's what happens mechanically. One partition serves the hot key. Its CPU pins, its request queue backs up, and latency climbs — not just for the hot key, but for every key that happens to live on that partition. Innocent-bystander users, sharded onto that same partition by pure hash luck, now see multi-second responses for their own boring profiles. Meanwhile the sibling partitions idle.
Now the trap: your fleet-average dashboard. Average CPU across 64 partitions where one is at 100% and the rest at 5%? About 6.5%. Green. Average latency, when most requests still hit cold partitions? Looks fine. But if that one hot partition serves 30% of live traffic (which is the whole point — it's hot), then 30% of your requests are slow, your p99 has exploded, and your p90 is next. Aggregate metrics are averages, and averages hide exactly the thing you need to see.
So detection rule one: keep per-partition metrics — QPS, CPU, and p99 latency for each partition, drawn as a heat map, not blended into an average. Detection rule two: know your top-K hottest keys at any moment. You can't afford to count every key exactly (billions of counters), but a count-min sketch — a small probabilistic counter that overcounts slightly but never undercounts, covered properly in chapter 2.10 — gives you a live "top 50 keys by traffic" list for a few megabytes of memory. That list is not just a debugging tool; two rungs of the mitigation ladder below need it as an input.
Discord stored trillions of chat messages in Cassandra, partitioned by (channel_id, time_bucket). Sound sharding — but Discord servers obey a power law too. When a giant community got busy, every reader and writer in it converged on the same channel partition, and Cassandra developed what Discord's engineers described as unbounded concurrency on the hot partition: queries piled up, latency grew, which made more queries pile up, cascading until JVM garbage-collection pauses forced painful node babysitting — on-call engineers were regularly firefighting the messages cluster. Their fix came in two layers. First, they built intermediary "data services" in Rust that do request coalescing: when a thousand users request the same message at once, the service issues one database query and fans the answer back out to all thousand waiters. Second, they migrated the whole thing from 177 Cassandra nodes to 72 ScyllaDB nodes in 2022–2023. Read p99 fell from 40–125ms to 15ms; write p99 from 5–70ms to 5ms. Notice the shape of the lesson: the partition key design was reasonable, and they still needed a traffic-shaping layer in front of the database, because no key schema makes a million people care about a million different things.
You can now detect the fire. Time to fight it — and the order in which you reach for the extinguishers is itself graded.
Learn these in order, cheapest first, and always say what each rung costs. Interviewers grade the ordering as much as the list: jumping to an exotic fix while skipping the cheap one reads as inexperience, and naming a fix without its price reads as blog-post knowledge.
For a read-hot key, the first answer is the boring one from chapter 1.6: cache it. Forty million people opening one profile don't need forty million database reads — the profile changes rarely. A cache with even a 5-second TTL absorbs virtually all of that heat: one database read per 5 seconds instead of thousands per second.
But run the numbers before declaring victory, because the hot key follows you into the cache. Your cache cluster is itself sharded by key — so all traffic for the hot key lands on one cache node. A single Redis node handles on the order of 100K ops/sec; a truly viral key can see over a million reads/sec. Congratulations: you moved the hot partition from the database into Redis. Same fire, more expensive building.
The senior finish is a local L1 cache: a small in-process cache (just a hash map with TTLs) inside each app server, holding only the few dozen hottest keys — fed by that top-K list from your detection layer. Now 200 app servers each answer from their own RAM, the Redis shard sees 200 refresh requests every few seconds instead of a million reads, and nobody melts. The cost: staleness of a few seconds, and 200 servers that may briefly disagree. For a follower count or a profile photo, nobody notices. For an account balance, this rung is not available — which is exactly the kind of caveat to say out loud.
If the read-hot data can't tolerate cache staleness, or the hot set is too big for L1 caches, replicate unevenly: give the hot partition five read replicas while cold partitions keep two. Chapter 1.8 gave you the machinery; the senior twist is applying it asymmetrically, driven by the heat map. Costs: replication lag (reads may be slightly behind the primary), extra hardware for one slice of the keyspace — and it does absolutely nothing for writes, because every replica must still apply every write. Which brings us to the harder half of the problem.
Now suppose the key is write-hot: a like counter on a viral post taking 8,000 increments per second, a live-match score, one IoT tenant firehosing telemetry at a single device row. Caching is useless — caches absorb reads, not writes. Replicas are worse than useless — they multiply the write load. A single partition sustains maybe 1,000–10,000 writes/sec depending on the store. You need the writes to land on more than one partition. But the key is the key — same key, same partition. That's the whole contract of sharding.
So break the contract deliberately. Salting: instead of writing to likes:post789, append a random suffix from 0 to N−1 and write to likes:post789#0 … likes:post789#15. Sixteen different keys hash to sixteen different partitions, and your 8,000 writes/sec become 500/sec per partition. Fire out.
Now pay the bill: to read the like count you must query all sixteen sub-keys and add them up — a scatter-gather read. One read became sixteen, and your read latency is now the latency of the slowest of the sixteen. Say this trade-off exactly this way in an interview: "I traded read cost for write relief." That one sentence is the difference between knowing the trick and understanding it.
Two refinements make this rung senior-grade. First: only salt the hot keys. If you salt every key with N=16, every read in your system becomes sixteen reads — you've bought fleet-wide 16x read amplification to fix a problem that maybe fifty keys have. Salt the keys on your dynamically tracked top-K list, and leave the billion cold keys alone. Second: pick N from arithmetic, not vibes. 8,000 writes/sec against a 1,000/sec-per-partition budget → N=10 with headroom, not N=100 "to be safe" — because every notch of N is a notch of permanent read cost on that key.
A bank assigns every customer to one teller — tidy, until payroll day, when a single giant company's deposits swamp teller #3 while the others file their nails. Salting is telling that one company: "your deposits can go to any of eight tellers." The queue vanishes. But now, to know the company's balance, the manager must walk to all eight tellers and add up eight ledgers — and the answer arrives only when the slowest teller finishes counting. Faster deposits, slower audits. Regular customers still use one teller, because making everyone's balance an eight-ledger scavenger hunt would be madness.
Salting is a patch you apply when a key catches fire. A compound key is the same idea applied at design time, using a real attribute instead of a random number. Don't partition messages by channel_id — partition by (channel_id, time_bucket), as Discord did with 10-day buckets, so one giant channel's history spreads over many partitions instead of growing one partition forever. Don't key IoT telemetry by device_id — key it by (device_id, hour). Don't key a multi-tenant table by tenant_id alone — key it by (tenant_id, entity_id) so the whale tenant's rows spread out.
The cost is the same scatter shape as salting, but bounded and meaningful: a query spanning three time buckets does three reads. And be honest about what compound keys don't fix — Discord's schema already had time buckets, and they still got hot partitions, because bucketing spreads a key's data over time, not this instant's traffic. When forty thousand people are reading the newest bucket right now, the newest bucket is hot. Compound keys prevent the chronic disease; they don't stop the acute attack.
Now the canonical exam question: the social feed. Chapter 1's push-vs-pull returns, at senior difficulty. Push (fan-out-on-write) precomputes every follower's feed at post time — reads are cheap cache hits. But a celebrity with 100M followers turns one tweet into 100M writes. Pull (fan-out-on-read) writes once and assembles feeds at read time — but now every feed load queries hundreds of followees. Each pure strategy dies at scale, just at opposite ends.
The senior answer refuses to pick one policy for all users: push for normal accounts, pull for celebrities, merge at read time. When a normal user posts, fan out to their 200 followers — trivial. When a celebrity posts, write it once to a celebrity-posts store and fan out nothing. When you open your feed, the system fetches your precomputed timeline and merges in recent posts from the handful of celebrities you follow. The load profile of each entity class gets the strategy that suits it, with a follower-count threshold deciding who counts as a celebrity.
In Raffi Krikorian's famous "Timelines at Scale" talk (he ran Twitter's platform engineering), Twitter had about 150 million active users and roughly 300K QPS of timeline reads, served by fan-out-on-write: every tweet was copied into each follower's Redis-backed home timeline. Then follower counts exploded. Lady Gaga had around 31 million followers — one tweet from her meant ~31 million timeline writes, and a fanout that could take minutes to finish. The visible symptom was almost comic: replies raced ahead of the original. A follower with a fast fanout path would receive someone's reply to Gaga's tweet before Gaga's tweet itself reached their timeline — conversations arriving before the thing they replied to. Twitter's fix was the hybrid: high-follower accounts are excluded from fanout entirely; their tweets are fetched and merged into timelines at read time, while ordinary accounts keep the push path. This design — push for the many, pull for the few, merge at read — is now the textbook answer, and interviewers expect you to reinvent it on the spot when they say "100 million followers."
Letters from your friends get delivered to your mailbox — push, done once per letter, cheap because your friends write to a handful of people. But the morning newspaper isn't hand-delivered to every reader by its columnist; you pick it up from the newsstand when you want it — pull, because one writer with a million readers can't visit a million mailboxes. Your morning reading merges both: mailbox first, newsstand on the way to work. Nobody thinks this is exotic; it's just matching the delivery method to the audience size.
Everything so far is you, the engineer, spotting heat and applying a fix. The final rung builds reflexes into the platform: load-aware partitioning, where the system itself watches per-partition heat and splits or relocates hot ranges automatically instead of trusting hash arithmetic; and request coalescing (Discord's trick from earlier — often called "singleflight"), where a thousand concurrent identical reads collapse into one database query whose answer is shared by all thousand waiters. Coalescing deserves special respect: it caps read amplification at the source, works even for cache misses, and costs almost nothing.
Amazon's DynamoDB enforces hard per-partition limits: 3,000 read capacity units and 1,000 write capacity units per second, per partition. For years, a hot key meant throttling errors even when the table as a whole had capacity to spare: the fleet bored, one partition drowning. DynamoDB's answer, rolled out through 2018–2019 and on by default for every table, is adaptive capacity: the system detects sustained heat on a partition and reacts on its own. It shifts throughput allowance toward the hot partition instantly, and if heat persists it performs a "split for heat" — dividing the hot partition into two, each on separate hardware. In the extreme case, AWS documents that a single relentlessly hot item can end up isolated in a partition of its own, getting the full 3,000/1,000 to itself. The celebrity literally gets a private server. The design lesson: even a database run by one of the best infrastructure teams on Earth cannot make the hot-key problem disappear — the per-item ceiling still exists — but a platform that detects and splits automatically turns a 3am page into a non-event. That's the end state your own architecture should aspire to.
One more grim scenario completes the picture: the partition is on fire right now, and the real fix is splitting it — which means moving live data off a machine that's already at 100%. Resharding under fire is a zero-downtime migration problem: dual writes, backfill, verified cutover, and the discipline to contain the heat with rungs 1 and 6 first so the migration itself doesn't tip the node over. Chapter 2.11 is that playbook; here, just know the sequence — contain first, split calmly second.
| Rung | Fixes | The price | Reach for it when |
|---|---|---|---|
| Cache + local L1 | Read heat | Seconds of staleness; brief cross-server disagreement | Read-hot, staleness-tolerant (profiles, counts) |
| Extra replicas on hot range | Read heat | Replication lag; hardware for one slice; writes untouched | Read-hot, can't tolerate cache staleness |
| Key salting (N ways) | Write heat | Every read of that key becomes N reads (scatter-gather) | Write-hot counters/streams; reads infrequent or batched |
| Compound keys | Chronic growth + heat | Multi-bucket queries; must be designed in early | Schema design time — always consider it |
| Hybrid per entity class | Fanout explosions | Two code paths; threshold tuning; merge logic at read | Power-law fanout: feeds, notifications |
| Adaptive splits + coalescing | Everything, automatically | Platform complexity; still a per-item ceiling | Platform-scale; recurring unpredictable heat |
Calibration, so you neither undershoot nor gold-plate. The bell that rings this topic is unmistakable: "one user has 100M followers," "a post goes viral," "one tenant is huge," or a suspiciously specific "traffic is uneven." When you hear it, an L5 answer volunteers — unprompted — this exact sequence:
That's the whole L5 bar: detection story, read/write classification, ordered mitigations with trade-offs stated, one committed recommendation. Ten sentences, maybe ninety seconds. What you do not need — staff-level garnish that wastes your clock if you lead with it: count-min sketch error bounds and internals (name the tool, cite chapter 2.10's territory, move on), designing a custom load-aware rebalancer from scratch, Zipf exponents, or DynamoDB's internal split heuristics. Mention that adaptive systems exist ("DynamoDB does this automatically — split for heat") as one sentence of awareness, not five minutes of trivia. A crisp ladder beats a probabilistic-data-structures lecture every single time.
How this topic is actually probed — usually as an ambush on a design you already drew:
Explain to a junior engineer, in four or five sentences, what a hot partition is, why perfect sharding doesn't prevent it, and the first two things you'd do about one.
Sharding spreads keys evenly across partitions, but it can't spread attention — and real traffic is wildly uneven, so when one key goes viral, the single partition holding it gets hammered while its siblings sit idle. Your average-based dashboards will look healthy the whole time, because one saturated partition out of sixty-four barely moves the average — you need per-partition metrics and a list of the hottest keys to even see it. First move: if the key is being read a lot, cache it — including a tiny cache inside each app server, because otherwise the hot key just melts one cache node instead of one database node. If the key is being written a lot, caching does nothing; instead you split the key into N sub-keys on different partitions, which spreads the writes but means every read now has to query all N pieces and merge them. That read-cost-for-write-relief trade is the heart of the whole topic.
A World Cup final. Your live "match hub" page has one likes counter taking 500K likes per minute at peak. Your database sustains about 1,000 writes/sec per partition comfortably. Size the salting: pick N, justify it, and state what reads of the counter now cost — plus one trick to make that read cost irrelevant to users.
500K/min ≈ 8,300 writes/sec. Against a 1,000/sec budget you need at least 9 partitions; N=16 gives ~520/sec per sub-key, roughly 2x headroom for spikes — justified, not vibes. (N=100 would work too but buys nothing except more read cost.) Reads now scatter-gather 16 sub-keys, latency = slowest of 16. The trick: nobody needs a like count that's fresh to the millisecond — run a small aggregator that sums the 16 sub-keys once per second and caches the total, so user reads hit one cached value and the scatter-gather happens once per second regardless of viewer count. Write relief, and the read bill is paid by one background job instead of millions of users.
Users report multi-second load times. Fleet dashboard: average CPU 10%, average latency 80ms, error rate normal. Name the three graphs you pull up next, in order, and what each would tell you.
(1) Per-partition heat map of CPU/QPS — one red cell among green confirms a hot partition and names it. (2) Per-partition p99 latency — confirms the red partition is where the slow requests live, and shows the blast radius (every key co-resident on it). (3) Top-K hottest keys (count-min-backed) — tells you which key and whether it's read-hot or write-hot, which selects the mitigation rung. Bonus senior point: check whether the slow user reports correlate with the hot partition's key range. The pattern to internalize: averages said "fine" three times; per-partition views found the fire in three graphs.
What breaks if, "to be safe," you salt every key in the system with N=20 instead of only the hot ones?
Every read in the system becomes 20 reads. Fleet-wide read amplification of 20x: your database fleet needs to serve 20x the query volume for the same user traffic, every read's latency becomes the slowest of 20 partitions (so p99 gets dramatically worse for everyone), and costs scale accordingly — all to fix a problem that perhaps fifty keys actually have, since heat is power-law by nature. This is why salting must be targeted: track the top-K hot keys dynamically, salt only those (a per-key flag or a hot-key config the writers consult), and unsalt when they cool. Defensive blanket salting is the over-engineering version of this topic, and interviewers probe for exactly this misjudgment.
Your B2B analytics SaaS has 4,000 tenants sharded by tenant_id. One whale tenant generates 40% of all traffic and their queries are slowing down neighbors on their partition. Sketch your response — immediate, structural, and organizational.
This is the celebrity problem wearing a suit: the whale is your Lady Gaga. Immediate: per-tenant rate limits so the whale can't starve co-tenants, plus caching/coalescing on their heaviest read queries. Structural: move the whale to dedicated partitions (or a dedicated cell, in chapter 2.5's language) — a live migration, dual-write then cutover per chapter 2.11 — and change the key schema to (tenant_id, entity_id) so any single tenant's data spreads over multiple partitions instead of concentrating. Organizational: whales are predictable, unlike viral posts — sales knows who the big customers are, so provision dedicated capacity at onboarding rather than discovering it from a heat map. Hybrid thinking again: normal tenants share infrastructure (push-model economics), whales get isolated treatment (pull-model economics), and a threshold decides who's who.
It's 11:55pm, five minutes before your flash sale. You did everything right. You load-tested at 3x normal traffic. You warmed the caches. Autoscaling is on. The war room has snacks. At midnight the sale opens and traffic climbs to 2.2x normal — comfortably under plan. Everyone relaxes.
At 12:04, your payment provider slows down. Not down — slow. Their p99 goes from 200 milliseconds to 2 seconds. Annoying, but survivable, right? At 12:09, your checkout service stops responding. At 12:14, product pages die — pages that never touch payments at all. By 12:19 every dashboard is red. Traffic is still only 2.2x normal.
The next morning, one graph from the postmortem explains everything: the requests hitting your backend peaked at 9x normal, while requests from actual users never passed 2.2x. Where did the other traffic come from? From inside the house. Your own retries. Nobody attacked you. You attacked yourself, and your system had no way to say "stop."
This chapter is the anatomy of that outage, and the toolkit that prevents it: timeouts, retries done right, circuit breakers, bounded queues, backpressure, and load shedding. At the senior bar this is not optional garnish. "What happens at 10x load?" is the single most reliable E5 probe, and your answer needs to be loaded before the question is asked.
Overload almost never kills a system directly. The system's reaction to overload kills it. The sequence is so repeatable it deserves to be memorized like a folk tale.
Act one: the queue grows. A server running at 80% capacity is fine. Push arrivals past 100% of what it can process and requests start waiting in line — in a queue, a thread pool backlog, a connection pool, somewhere. Here's the brutal arithmetic of queues: your waiting time is roughly the queue depth divided by the processing rate. A queue 10,000 deep in front of a service that handles 1,000 requests per second means every new arrival waits 10 seconds before any work starts.
Act two: clients give up, but the work doesn't. No user waits 10 seconds. Clients time out and disconnect — but their requests are still sitting in the queue. The server faithfully processes them and sends answers to callers who left long ago. Call this zombie work: real CPU spent producing responses nobody will receive. Past a certain queue depth, your server can be 100% busy and 0% useful.
Act three: the retries arrive. Every client that timed out tries again. If clients retry 3 times, each original request becomes up to 4 attempts. Your load just quadrupled — at the exact moment your capacity is at its worst. This is the cruelest property of naive retries: they cost nothing when the system is healthy and multiply load precisely when it is sick.
Act four: the starvation spreads. Meanwhile, in every service that calls the slow dependency, threads are parked waiting for answers. Thread pools are finite — say 200 threads per server. When all 200 are stuck waiting on a slow payment provider, that server cannot serve anything, including the product pages that never touch payments. This is why "one dependency got slow" becomes "the whole site is down": the slowness travels through shared thread pools, connection pools, and memory.
Act five: nothing can recover. Even when the original fault heals, the accumulated queues and swarming retries keep every service pinned at maximum load. The outage now sustains itself. Someone has to break the loop by hand.
Every tool in this chapter attacks one arrow in that diagram. Timeouts cap how long threads are held hostage. Smart retries shrink the feedback arrow. Circuit breakers and bulkheads stop the starvation from spreading. Bounded queues kill zombie work. Load shedding caps what enters.
Start with the multiplication, because the numbers are shocking. Suppose your stack is four layers deep — edge, service A, service B, database — and every layer retries failed calls 3 times. One user request that fails at the bottom can generate 4 × 4 × 4 = 64 attempts against the database. You built a 64x load amplifier and armed it to trigger exactly when the database is struggling. Amazon's Builders' Library is blunt about it: retries are "selfish" — each retry spends more of the server's scarce capacity to improve one caller's odds — so at scale, your biggest DDoS threat is your own retrying client fleet.
Three guards turn retries from a weapon into a tool.
Guard one: exponential backoff — with jitter. Waiting longer between attempts (1s, 2s, 4s, 8s) helps, but it has a sneaky flaw: every client that failed at the same moment retries at the same moment. Backoff without randomness turns a crowd into a synchronized army that attacks in waves — the waves just get further apart. The fix is jitter: randomize each delay, for example sleep = random(0, base × 2^attempt). Now the herd arrives as a gentle drizzle spread across time instead of a battering ram. In an interview, saying "exponential backoff" without "jitter" is a half-answer, and interviewers know it.
Guard two: retry budgets. Backoff spreads retries out; it doesn't limit how many there are. A retry budget does: cap retries at a percentage of total traffic — say, retries may never exceed 10% of requests. While the dependency is healthy, nothing changes. When it breaks, amplification is capped at 1.1x instead of 4x, and the excess failures return errors immediately. A simple token bucket implements it.
Guard three: only retry what's safe, and only in one place. Retrying a read is free. Retrying "charge the card" without an idempotency key (p2c02) is how customers get billed twice — the timeout tells you the outcome is unknown, not that it failed. And pick one layer of the stack to own retries (usually near the edge, where you know the user is still waiting); inner layers should fail upward, not multiply. That single decision turns 4 × 4 × 4 back into 4.
Power companies fear the moment after a blackout more than the blackout itself. When electricity returns, every fridge, water heater, and air conditioner in the city starts compressors at the same instant — demand surges far above normal and trips the grid again. Engineers call it cold load pickup, and the cure is to re-energize neighborhood by neighborhood, staggered in time. That's your retry storm: the original fault is brief, but the synchronized reconnection is what keeps knocking the system back down. Jitter is the staggering; load shedding is choosing which neighborhoods wait. The analogy is honest on one more point: the grid solves it with a human plan made in advance — and so should you.
Early on a Sunday morning, a brief network disruption hit DynamoDB's storage fleet in US-East. Storage servers periodically confirm their "membership" — which table partitions they own — with an internal metadata service. The recently launched Global Secondary Indexes feature had quietly made those membership lists much larger, pushing responses close to the timeout limit. When the network blip ended, a huge number of storage servers all requested their membership at once. The oversized responses timed out; servers that couldn't confirm membership took themselves out of service and kept retrying — flooding the metadata service so completely that even healthy requests failed. The network fault lasted moments; the retry storm held the door shut for hours. AWS engineers had to explicitly pause requests to the metadata service — shedding load on their own system — just to get enough breathing room to add capacity. Error rates didn't return to normal until after 7am, and the blast radius included SQS, CloudWatch, and EC2 Auto Scaling. The public postmortem reads like this chapter's checklist: they increased metadata capacity, lengthened timeouts, and reduced retry aggressiveness. The lesson seniors quote: the fault triggers the outage, but retries are what sustain it.
A remote call without a timeout is a thread donated to the void. Many HTTP and database client libraries default to timeouts of minutes — or infinity. During the cascade's act four, those defaults decide how fast your thread pools drain: with a 30-second default, one slow dependency eats 200 threads in seconds.
Two rules make timeouts a system instead of a superstition.
Rule one: derive timeouts from data, not vibes. Set a call's timeout from its observed latency — a common choice is around p99 plus margin. Too tight and you fail healthy-but-slow requests, triggering needless retries; too loose and you hold threads hostage — a 10-second timeout on a 50ms call is a 200x hostage multiplier, not "safety."
Rule two: budget timeouts end-to-end. The caller's timeout must be longer than everything the callee might do — all attempts plus backoff — or you've built a machine for manufacturing zombie work. If the edge gives up after 1 second but service A's plan is "two tries at 700ms each," then A routinely burns 1.4 seconds computing answers for a caller who left at 1.0s — and the edge's retry arrives while the first attempt is still running. Sum it, write it down, check the inequality: caller timeout > (tries × per-try timeout) + backoff + own work. Mature systems propagate the remaining budget with the request — gRPC calls this a deadline — so a callee with 300ms left doesn't start a 500ms operation. Naming deadline propagation is a strong senior move; implementing the plumbing is not expected at the whiteboard.
Timeouts cap the damage per call, but they still pay full price per call. If the payment provider is going to be down for five minutes, why spend 800ms and a thread on every request to rediscover that fact ten thousand times? Once you know a dependency is sick, the right response is to stop calling it.
That's a circuit breaker: a small state machine wrapped around each dependency.
Failing fast is only half the value. The other half is the fallback: what you return while the breaker is open. Options, in rough order of preference: a cached or stale copy of the data; a sensible default (an empty "recommended for you" row); a degraded response ("payments are busy, your order is saved and will process shortly"); or an honest, instant error. Even the bare error is a win — it answers in microseconds instead of burning a thread for 800ms, which keeps your thread pool alive and gives the sick dependency silence in which to recover.
Which brings us back to act four: why did product pages die when only payments were slow? Because every request, payment-related or not, drew threads from the same pool. The fix is bulkheads: give each dependency its own small, isolated resource pool — its own threads, connections, or a semaphore cap on concurrent calls. If payments gets 20 threads out of 200, then when payments hangs, it exhausts its 20 and the other 180 keep serving product pages. The name comes from ships: watertight compartments so one breach floods one section, not the hull.
The name "circuit breaker" is borrowed from your house's electrical panel, and the borrow is honest with one twist. The breaker in your wall exists because the alternative — wires overheating inside your walls until they ignite — is catastrophically worse than the inconvenience of a dark kitchen. It converts "invisible slow-burning disaster" into "visible instant failure," which is exactly what the software version does: an open breaker turns a thread-eating slow death into a clean, fast, obvious error. The twist: your wall breaker protects only your house, and a human must walk over and reset it. The software breaker protects both sides — the caller's threads and the struggling dependency's remaining capacity — and half-open is its built-in electrician, probing on a timer to see whether it's safe to reset.
Around 2011, Netflix's API layer fanned out to dozens of backend services: with 30 dependencies each at 99.99% uptime, something is always failing somewhere — and any one dependency going slow could exhaust the API tier's threads and take down the whole product. After exactly such cascades, Netflix built Hystrix and open-sourced it in 2012: circuit breakers, per-dependency thread-pool bulkheads, timeouts, fallbacks, and a live dashboard, wrapped around every remote call. It became the defining resilience library of the microservices era. Then in November 2018, Netflix put Hystrix into maintenance mode — not because the ideas failed, but because one of them aged badly: Hystrix's thresholds and pool sizes were static configuration, hand-tuned numbers that drifted out of date every time traffic patterns changed. Netflix's newer direction is adaptive concurrency limits: instead of a human guessing "20 threads for payments," the client continuously measures latency and adjusts its own concurrency cap, the way TCP discovers available bandwidth. The community successor for the classic pattern is resilience4j, which Spring officially recommended as Hystrix's replacement. The senior takeaway: breakers and bulkheads are permanent ideas; hand-tuned static thresholds are the part that rots.
Part 1 taught queues as shock absorbers: buffer a burst, work it off when the spike passes. That's true for bursts. Under sustained overload — arrivals persistently above capacity — a queue doesn't absorb anything. It just grows, and with it grows latency (depth ÷ throughput), memory usage, and the fraction of the queue that is zombie work whose callers have already gone. An unbounded queue converts "reject some users now" into "crash for everyone later, after wasting maximum effort." That is not a safety feature. It's a deferred outage with interest.
So: every queue in your design gets a size and a full-policy, chosen out loud. When the queue is full, you have exactly three options, and the choice is a design decision, not an accident:
When demand exceeds capacity, somebody, somewhere, must either slow down or be turned away. The senior framing is that these are two named mechanisms applied at two different places:
Backpressure operates inside your boundary: an overloaded component signals its upstream to slow down, and the signal propagates until it reaches something that can safely absorb it. A bounded queue that blocks its producer is backpressure. TCP flow control is backpressure at the internet's foundations: receivers advertise what they can accept. Kafka consumers pull messages rather than having them pushed, so a slow consumer naturally reads slower — pull systems get backpressure for free. The key property: inside your walls, you control both ends of every arrow, so "please slow down" is enforceable.
Load shedding operates at the edge: you cannot make the public slow down, so you reject excess work before it costs you anything. A cheap, instant 429 Too Many Requests with a Retry-After hint, served at the load balancer or gateway. Two rules make shedding effective. First, reject early and cheaply — shedding a request after you've done auth, session loading, and two database reads saves almost nothing; the whole point is that a rejection costs microseconds while a served request costs milliseconds. Second, shed by priority, not at random: classify requests — checkout and login are sacred, browsing is important, analytics beacons and speculative prefetches are disposable — and drop from the bottom. Dropping 100% of analytics is invisible; dropping 5% of checkouts is a revenue incident. The priority list must be written before the incident, because 2am during an outage is the worst possible moment to debate whether courier location pings outrank review images.
The trigger for shedding is usually utilization-based admission: watch a signal that tracks saturation — CPU, queue depth, or in-flight request count — and when it crosses a threshold, start rejecting the lowest priority class, escalating to higher classes only as needed. Google's SRE guidance lands on the same stance: a service should serve at its true capacity and reject the excess cheaply and predictably — degraded service beats no service, and predictable rejection beats unpredictable collapse.
Notice how the two mechanisms connect: backpressure travels upstream hop by hop until it reaches the edge, and the edge — the one place that can't tell its upstream to slow down — converts the pressure into shedding. A system with backpressure but no shedding jams up at the front door; a system with shedding but no backpressure protects the door while the kitchen drowns. You need both, and saying that sentence at the whiteboard is a senior signal all by itself.
During the final step of a schema migration — an ordinary table rename — a significant portion of GitHub's MySQL read replicas hit a semaphore deadlock and crashed into recovery. The remaining healthy replicas inherited the full production read load, buckled, and crashed too. Then the recovery itself failed: as each replica finished crash-recovery and rejoined, the full weight of pent-up production traffic slammed into it immediately, and it would — in GitHub's own words — "temporarily recover from their crash-recovery state only to crash again due to load." A classic recovery crash-loop: the herd re-killed every node that stood up. GitHub promoted every healthy internal replica into production to add capacity; it wasn't enough. What finally worked was deliberate load shedding: proactively holding production traffic away from recovering replicas until they had fully finished the rename, accepting deeper degradation now to make recovery possible at all. The incident degraded Actions, API, Issues, Pull Requests, and webhooks for about 2 hours 50 minutes; the remediation included functional partitioning — bulkheads at the database layer — so one cluster's failure can't take the whole product down. The transferable lesson: recovery is itself an overload event. If your design can't shed load, it can't come back up, because returning traffic re-kills each node the moment it stands.
Shedding answers "who gets rejected." Degradation answers a gentler question: "what can we serve instead?" The options form a menu, and strong teams write it down before the incident, with a feature flag next to each item:
These menu items are exactly what your circuit-breaker fallbacks return. Breakers decide when to degrade; the menu decides what degraded looks like. Designed together, users experience a slightly duller product; designed never, users experience a 500.
This topic surfaces in one costume: "Nice design. What happens at 10x load?" At E5, that question is not an invitation to think — it's a check that you already have. The passing answer is a preloaded, ordered narrative: "The first thing to saturate is X — here's the number that says so. The queue in front of it is bounded at N, so when it fills we reject rather than balloon latency. The edge starts shedding in this priority order; these three features degrade to cached fallbacks via their breakers; checkout stays up. Here's the alert that fires before users notice." Failure modes volunteered, numbers attached, priorities pre-decided.
Calibration — what L5 must volunteer, unprompted, versus what is staff-level garnish:
The trap to avoid is the vocabulary answer: "I'd add circuit breakers and rate limiting" with no thresholds, no fallbacks, no priority order. Interviewers hear that the way you'd hear "I'd add security." The mechanisms only count when they come with numbers and decisions attached.
How this actually gets probed at E5/L5 — usually as a live stress-test of the design you just drew:
A passing answer commits: "Retries live at the edge only: two attempts, full jitter, 10% budget, idempotency keys on writes. Payments gets a breaker at 50% errors over 10 seconds with a queue-and-confirm fallback, and its own connection pool so checkout being sick never touches browsing."
Explain to a junior engineer, in five or six sentences, why adding retries can take down a healthy system — and what three guards make retries safe.
When a service gets slow, every caller that times out tries again — so a service that's already drowning suddenly gets two, three, four times its normal traffic, exactly when it can least handle it. Worse, all those retries tend to arrive at the same moment, like everyone redialing the instant a call drops, so the load comes in synchronized waves that knock the service down each time it tries to stand up. Three guards fix this. First, exponential backoff with jitter: wait longer between attempts and randomize the wait, so the herd spreads out instead of stampeding. Second, a retry budget: retries may never exceed some small fraction of total traffic — say 10% — so amplification is capped no matter how bad things get. Third, only retry operations that are safe to repeat: a timed-out "charge the card" didn't necessarily fail, so writes need idempotency keys before they're retryable, and only one layer of the stack should retry at all — otherwise three layers of three retries each turns one request into dozens.
Your stack is edge → service A → service B → database, and every layer independently retries failures up to 3 times (so up to 4 attempts each). The database has a rough moment. What's the worst-case number of database attempts generated by one user request? What does it become if only the edge retries? And with a 10% retry budget at every layer instead?
Worst case: 4 × 4 × 4 = 64 database attempts from one user click — each of the edge's 4 attempts drives A to 4 attempts, each of which drives B to 4. If only the edge retries and inner layers fail upward: 4 attempts total, a 16x improvement from one policy decision. With a 10% budget per layer, amplification is capped at 1.1 per layer ≈ 1.33x overall — the system stays roughly at normal load even during total dependency failure. This is why "retries in one place, budgeted" is the senior answer, and why every layer having its own generous retry loop is a 64x amplifier hiding in plain sight.
Code review: the edge calls service A with a 1-second timeout. Service A calls service B with a per-try timeout of 700ms and 2 tries, plus ~100ms of its own work. Find the bug and propose fixed numbers.
A's worst case is 700 + 700 + 100 = 1500ms, but its caller gives up at 1000ms. So whenever B's first try times out, A's second try is pure zombie work — and worse, the edge's own retry of the whole request arrives while A is still mid-flight, doubling load during slowness. Fix by making the inequality hold: e.g., edge timeout 1s → A gets a ~900ms budget → B gets 2 tries at 350ms each plus ~100ms backoff and 100ms own work ≈ 900ms. (Or keep 700ms per try but allow only one try.) The senior habit: walk every hop and check caller timeout > tries × per-try + backoff + work — this bug ships constantly because each number looks reasonable alone.
A service processes 2,000 requests/second. A spike pushes arrivals to 3,000/second, and there's an unbounded queue in front. Clients time out at 5 seconds. How fast does the queue grow, when does the service stop doing any useful work, and what bound plus full-policy would you set?
The queue grows at 1,000 requests/second (arrivals minus capacity). Waiting time is depth ÷ 2,000/s, so once depth passes 10,000 (after just 10 seconds of overload), every queued request waits over 5 seconds — beyond the client timeout. From that moment the service runs at 100% capacity producing answers nobody will receive: fully busy, zero useful. A sane bound is capacity × the client timeout with margin — e.g., 2,000/s × 2s = 4,000, keeping worst-case queue wait around 2 seconds. Full-policy: reject newest with an instant 429 (callers get fast honest errors while their wait would have been useless anyway). The one-liner worth memorizing: a queue deeper than capacity × client-timeout is a machine for wasting your own CPU.
You run a food-delivery app. A regional traffic spike forces shedding. Rank these five request types from first-shed to never-shed, and justify: (a) analytics events, (b) restaurant browsing, (c) order placement, (d) courier location pings for active deliveries, (e) review photo loads.
One defensible ranking: shed (a) analytics first — zero user impact, often 30%+ of traffic; then (e) review photos — pages work without them; then degrade (b) browsing — serve cached restaurant lists, drop personalization, but keep it alive since browsing feeds ordering; protect (d) courier pings hard — they look like telemetry but they're the live product for every active delivery, and dropping them breaks tracking and dispatch (the trap in this question: "pings" sound disposable and aren't); never shed (c) order placement — it's the revenue event. The meta-lesson: rank by business damage per rejected request, not by how technical the traffic sounds — and write the ranking down before the incident, because you will not have this debate calmly at 2am.
It's 2:47am. Your phone goes off. The alert says: "CPU above 80% on web-14." You drag the laptop open, squint at a graph, and by the time it loads, CPU is back to 60%. Nothing else looks wrong. You go back to sleep, annoyed.
At 9am you learn the truth: checkout had been failing for 40% of users since 9pm. A bad deploy to the payment service. Eleven hours. Nobody was paged, because there was no alert on "users can't pay." There was an alert on CPU. The company lost a night of revenue, and the on-call engineer — you — was woken up once, for the wrong thing, and told everything was fine.
Every chapter so far taught you to build the machine: load balancers, caches, replicas, shards, queues. This chapter is the instrument panel. It's the least glamorous topic in the book and the one interviewers now use to separate people who have operated systems from people who have read about them. The rule is brutal and simple: you can't fix what you can't see — and worse, you can't even know it's broken.
"Observability" gets thrown around as a buzzword. Strip it down and it's three kinds of telemetry — data your system emits about itself — and each kind has a distinct job. Senior candidates don't recite the three names; they know which one to reach for and why, because each sits at a different point on a cost-versus-detail trade-off.
Metrics are numbers over time, pre-aggregated: requests per second, error rate, p99 latency, queue depth. A metric doesn't remember individual requests — it remembers counts and distributions. That's what makes metrics cheap: a million requests per second collapses into a handful of numbers per second. Cheap means you can keep them for months and query them in milliseconds, which is why metrics are the fuel for dashboards and alerts. Their job: detect that something is wrong, fast.
Logs are the opposite trade. A log line is one event, in full detail: this request, this user, this error message, this stack trace. Detail makes logs expensive — at scale, log storage is often a top-three infrastructure bill — and slow to search. Their job: explain what happened, after metrics told you something did. You don't watch logs to find out you're down; you dig through logs to find out why.
Traces answer the question neither of the others can: what happened to one request as it crossed ten services? At the front door, the request gets a unique trace ID. Every service that touches the request passes the ID along in a header — this hand-off is called context propagation — and records a span: "I, the payment service, spent 1,840ms on this request, 1,790ms of it waiting for the fraud check." Stitch the spans together and you get a waterfall showing exactly where the time went. Their job: locate the slow or broken hop in a distributed system, where "the request was slow" has ten suspects.
A hospital runs on the same three pillars. Metrics are the vitals monitors at the nurses' station — heart rate, oxygen, one glance covers forty patients, and an alarm sounds the moment a number crosses a line. Logs are the patient charts — every dose, every observation, written down in full; nobody reads charts all day, but when something goes wrong, the chart is where the answer lives. Traces are the wristband ID that follows one patient from admission to radiology to surgery — scan it anywhere and you can reconstruct that one patient's entire journey and see exactly where they waited three hours. The analogy is honest about cost, too: monitors are cheap to watch, charts are expensive to write and read, and the wristband only works if every department scans it — miss one hand-off and the journey has a hole. That's context propagation.
Here's the mistake that takes down monitoring systems: treating cheap metrics like detailed logs. Metrics systems store one time series — one stream of numbers — per unique combination of labels. A metric labeled by endpoint (50 values), status (10), and region (5) creates 2,500 series. Fine. Now some well-meaning engineer adds a user_id label to "see which users get errors." Ten million users just turned 2,500 series into 25 billion. This is a cardinality explosion — cardinality means the number of distinct label values — and it will eat your monitoring cluster's memory and either crash it or bankrupt you, whichever comes first.
The rule: metrics get bounded, low-cardinality labels (status, region, endpoint template like /users/{id} — never the raw URL). Anything unbounded — user IDs, request IDs, email addresses — belongs in logs and traces, which store individual events anyway and are built for needle-in-haystack search. If you catch yourself wanting "a graph per user," you actually want a log query.
Given infinite things you could measure, what goes on the one dashboard everyone looks at first? Google's SRE book gives the answer that has become industry default — the four golden signals, and if you can only measure four things about a user-facing service, measure these:
| Signal | What it is | The question it answers |
|---|---|---|
| Latency | How long requests take — tracked separately for successes and failures (fast errors pollute the numbers) | Is it slow? |
| Traffic | Demand: requests/sec, messages/sec, streams open | How much is being asked of us? Did demand just double — or drop to zero? |
| Errors | Rate of failed requests — explicit (500s), or implicit (wrong content, 200-OK-but-empty) | Is it broken? |
| Saturation | How full the system is: memory, connection pools, queue depth, disk — the resource that runs out first | How close to the cliff are we? |
Notice traffic dropping to zero is a signal too — an error rate of 0% is not comforting when requests/sec is also 0 because your load balancer is sending traffic into a wall. And saturation is the forward-looking one: latency, traffic, and errors tell you about now; the queue that's 85% full tells you about twenty minutes from now. That connects straight back to backpressure in p2c07 — saturation metrics are how you see the pressure building before it becomes load shedding.
Every latency number in this book has come with a p-prefix — p50, p95, p99 — and this is the chapter that pays that debt. Take 100 requests: 99 finish in 100ms, one takes 10 seconds. The average is 199ms. Sounds healthy. But nobody experienced 199ms — 99 people had a fast experience and one had an awful one, and the average describes a user who doesn't exist. Averages smear the pain until it's invisible.
Percentiles keep the shape of the truth. Sort all response times: p50 (the median) is the time the middle request took — half were faster. p95 — 95% were faster, 5% slower. p99 — the threshold your slowest 1% cross. In our example, p50 is 100ms and p99 pierces into seconds, which is the real story. Interviewers expect senior candidates to talk in percentiles by reflex, and to know why the tail matters commercially: your heaviest users — the ones with the most photos, the biggest carts, the fullest inboxes — live disproportionately in the tail, because their requests touch the most data.
Two traps inside percentiles, both interview favorites:
Trap one: you cannot average percentiles. Server A has a p99 of 100ms, server B has a p99 of 500ms. The fleet p99 is not 300ms — it depends entirely on how much traffic each got and the shape of each distribution; the true answer could be almost anywhere between 100 and 500. A p99 is a property of one specific pile of numbers, and you can't combine two piles by averaging their summaries — you have to merge the piles. Real systems ship histograms (counts in latency buckets) from every server and merge those, or use clever compressed summaries called sketches. How those sketches work is a genuinely fun problem — it's a core deep-dive in p3c18 when you build a metrics system yourself. For now, the senior-level sentence is: "I'd export histograms per server and compute percentiles after merging, because averaging p99s is meaningless."
Trap two: tail latency amplification. Your service is admirably fast: 99% of calls under 10ms. Then you build a feed page that fans out to 100 backend calls — one per followed user, say — and waits for all of them. What fraction of page loads hit at least one slow call? One minus 0.99 to the 100th power: about 63%. Read that again. Each backend is slow 1% of the time, and most of your users experience the slow path, because the page is only as fast as its slowest call. Fan-out promotes your p99 into everybody's median. This single piece of arithmetic is why big systems obsess over tails: at Meta or Google scale, one user action fans out to hundreds of services, and a 1-in-100 hiccup somewhere is a certainty everywhere.
So you can see the system. Metrics detect, logs explain, traces locate, percentiles tell the truth. Now the harder question — the one that starts arguments between engineering and product: how good does the system have to be? "As reliable as possible" is not an answer. It's the absence of one.
Three acronyms, one chain, and interviewers love asking candidates to untangle them because the difference is exactly the kind of precision the job requires.
An SLI — Service Level Indicator — is a measurement: a carefully chosen ratio of good events to total events. "The fraction of checkout requests that succeed within 2 seconds, measured at the load balancer." Not CPU, not memory — those are causes, not experiences. A good SLI is the closest number you can compute to "was the user happy?"
An SLO — Service Level Objective — is your internal target for that SLI: "99.9% of checkout requests succeed within 2 seconds, over a rolling 30 days." Do the math on what 99.9% actually permits, because the math is the whole point. A 30-day month has 43,200 minutes. The 0.1% you're allowed to fail is 43 minutes of full outage per month. Every nine you add divides that by ten:
| Target | Allowed downtime / 30 days | What it implies |
|---|---|---|
| 99% | ~7.2 hours | Fix it during business hours |
| 99.9% | ~43 minutes | On-call with a pager; solid single-region design |
| 99.99% | ~4.3 minutes | Humans can't respond that fast — recovery must be automated; multi-zone at least |
| 99.999% | ~26 seconds | Multi-region failover, no maintenance windows, everything automated, enormous cost |
Stare at the 99.99% row. Four minutes is less time than it takes a human to wake up, open a laptop, and read the alert. Above three nines, you're no longer buying "more careful engineers" — you're buying automated failover, redundant everything, and progressively exotic architecture. Each extra nine multiplies cost roughly tenfold, and this is the trap interviewers set on purpose: ask a candidate what availability an internal analytics tool needs, and the junior answer is "five nines, reliability matters!" The senior answer is "99.5% — it's used by forty analysts on weekdays; a half-day outage is an inconvenience, and every nine past that is money stolen from features nobody voted to steal." Naming the right number, and defending why it isn't higher, is the strongest signal this chapter can teach you.
An SLA — Service Level Agreement — is an external contract: the promise in the customer's terms of service, with penalties (usually service credits, sometimes real money) when you miss it. Because breaching it costs money and trust, your SLA is always looser than your SLO: promise 99.9% to customers, hold yourself to 99.95% internally, and your own alarms fire while there's still runway before the contract burns. SLO is the tripwire; SLA is the cliff.
Think of a food delivery promise. The SLI is the measurement: what fraction of orders arrived within 30 minutes this month. The SLO is the kitchen's internal bar: "we hit 30 minutes for 97% of orders, or managers start asking questions." The SLA is what's printed on the app: "late orders over 45 minutes get a voucher" — an external promise, deliberately softer than the internal bar, with a penalty attached. The kitchen aims at 30 so it never pays out at 45. If the only number you track is the printed promise, you find out you're failing at the exact moment it starts costing money.
Here's the quiet genius hiding in an SLO. If your target is 99.9%, then 0.1% of failure is not just tolerated — it's yours to spend. That 0.1% is your error budget: 43 minutes a month of allowed badness. And the moment you treat it as a budget, it becomes a governance tool — a way to settle, with arithmetic, the oldest fight in software: developers want to ship fast (change causes outages), operators want stability (their pager rings). Both sides are right, and without a number they fight about it politically, forever.
The error budget replaces the fight with a rule agreed in advance: budget remaining → ship features, take risks, deploy on Friday if you like. Budget exhausted → feature launches pause, and engineering effort shifts to reliability work — better tests, safer rollouts, fixing the flaky dependency — until the budget recovers. No villain, no negotiation-by-seniority. The number decides.
Google created the Site Reliability Engineering role in 2003 under Ben Treynor Sloss, and inherited the classic war: product teams pushed launches, SRE teams — who carried the pagers — pushed back. Escalations were settled by whoever argued better, which scaled terribly. The fix, described in Google's 2016 SRE book, starts from an observation Treynor made famous: 100% is the wrong reliability target for basically everything. A user on a phone network that drops 1% of packets cannot tell the difference between 99.99% and 100% — the last fraction of a nine vanishes into the noise of everything between you and them, so the enormous cost of chasing it buys nothing anyone can perceive. So each service got an SLO agreed by both sides, and 100% minus the SLO became a shared budget. Budget left: SRE stands aside and launches flow. Budget gone: launches pause automatically, and product engineers redirect to reliability — a consequence both sides signed up for before anyone was angry. The same book records the funniest corollary: Google's Chubby lock service was so reliably above its SLO that other teams started assuming it could never fail — so SRE began deliberately taking Chubby down to burn the excess budget and flush out the teams that had quietly built on a promise nobody had made.
The budget also fixes alerting. The naive alert — "page me if error rate exceeds 0.1%" — fires on every 30-second hiccup, wakes a human, and by the time they've logged in, the blip is gone. Do that nightly and you've manufactured alert fatigue: the on-call learns that pages are noise, starts sleeping through them, and then misses the real one. The senior alternative is to alert on burn rate — how fast you're spending budget relative to the month you have. Errors at exactly 0.1% is burn rate 1: you'd land the month precisely at budget, no alarm needed. A total outage burns at 1000x. The practical scheme: page a human when the burn is fast (say, a rate that would eat a meaningful chunk of the monthly budget within hours — that's a fire), and file a ticket when the burn is slow but sustained (a rate that would exhaust the budget in a week or two — that's rot, fixable at 10am). A 30-second blip at burn rate 2 moves neither needle, and nobody wakes up. Alert volume drops by an order of magnitude, and every remaining page means real, budget-threatening user pain.
Alerting has a philosophy, and it fits in two lines. Page on symptoms, not causes. Every page must be actionable.
A symptom is user pain: checkout success rate dropping, feed p99 doubling, messages not delivering. A cause is machinery: CPU at 85%, one pod restarting, disk at 70%. The story that opened this chapter is the philosophy violated in both directions at once — noise with no action to take, and real pain with no page at all. Causes belong on dashboards, where you consult them after a symptom fires, to diagnose. Symptoms belong on pagers. The test for every alert you keep: when this fires at 3am, is there a specific action the on-call takes? If the honest answer is "look at it and go back to bed," delete the page or demote it to a ticket. An alert nobody acts on isn't an alert — it's a lullaby.
On August 1, 2012, Knight Capital — then handling a sizable share of all US stock trading — deployed new order-routing code to its eight SMARS trading servers. A technician missed one server. That eighth machine still ran a piece of code called Power Peg, dead since 2003, and the new deployment reused an old feature flag that switched the zombie code back on. When markets opened at 9:30am, the eighth server began firing millions of unintended orders into the market. It took 45 minutes to stop, moved prices across roughly 150 stocks, and left Knight with about $460 million in losses — the firm collapsed into a rescue acquisition within a year. The observability detail is the painful part, documented in the SEC's order: starting at 8:01am — 89 minutes before the market opened — an internal system sent 97 automated emails to Knight staff flagging an error, "Power Peg disabled," on the exact server that was about to malfunction. Nobody acted, and the SEC noted why: the emails weren't designed as alerts, just informational messages nobody was assigned to read. Ninety-seven warnings, zero pages, no kill switch, and a 17-year-old company gone in 45 minutes. This is what "every alert must be actionable, and someone must own it" costs when it's violated: an unowned, unactionable alert is indistinguishable from silence.
One more habit separates operators from tourists: monitor from outside, too. Your own dashboards share fate with your infrastructure — during the big S3 outage of 2017, Amazon couldn't turn its own status dashboard red because the dashboard's icons were stored in S3. A trivial external probe that hits your public endpoint from somewhere else answers the only question that ultimately matters: can users, in the real world, reach you? If your telemetry pipeline dies with your service, your dashboards will report serenity all the way down.
Back in p0c01 you learned the thirty-second ops closing: say what you'd monitor and how you'd ship. This chapter gives that closing its machinery, because deploys cause most outages — the system was fine, then you changed it. So the last piece is wiring the SLO into the deployment path itself.
A canary release sends the new version to a small slice first — 1% of traffic, or one region — while the old version serves everyone else, exactly like the miner's canary that met trouble before the miners did. Now the golden signals earn their keep: compare the canary's latency, error rate, and saturation against the baseline fleet, side by side, same time window. Clean after an hour? Promote to 25%, then 100%. Degraded? Automatic rollback — no human decision required, because the trigger was agreed in advance: the canary breached the SLO, so it dies. Pair this with feature flags — switches that turn features on per-user-slice without redeploying — and risky changes get two independent safety nets: flags decouple releasing a feature from deploying code, so rollback of a bad feature is flipping a flag in seconds, not shipping binaries. The loop closes neatly: SLIs measure, SLOs judge, and the rollout system acts on the judgment before your error budget — rather than your users' patience — takes the hit.
Calibration matters, because this topic has a sharp line between senior table stakes and staff-level garnish. Here's the line.
L5 must volunteer, unprompted, for their own design: two or three named SLIs ("checkout success rate within 2s, measured at the LB; message delivery p99"), one concrete SLO with the downtime math said out loud ("99.9% — that's 43 minutes a month, which one on-call rotation can honestly defend"), what pages a human and why it's a symptom ("delivery success below target over 10 minutes pages; CPU never pages"), and a rollout answer ("canary 1%, compare golden signals, auto-rollback on SLO breach"). That's four sentences, maybe thirty seconds, and it's the difference between "mentions monitoring" and "has been on call." Saying the word "monitoring" without an object is a checkbox; naming the SLI is the signal. Equally senior: refusing nines — "this admin tool gets 99.5%, and here's why more would be waste."
Staff-only garnish (skip unless asked): the internals of percentile sketches (that's p3c18's problem to build), exact multi-window burn-rate thresholds, OpenTelemetry propagation formats, exemplars linking metrics to traces, tracing sampling strategies. Knowing these exists is fine; deep-diving them in an unrelated design burns minutes you owe to the crux. If the role is observability infrastructure itself, the calibration flips — but then you're in p3c18, designing the metrics system as the problem.
This topic is rarely its own question at L5 — it's an ambush inside every question. The interviewer waits for your design to settle and then probes whether you'd survive operating it:
Explain SLOs and error budgets to a junior in five or six sentences — including why a team would ever want less than 100% reliability.
An SLO is a target you set for a number that tracks user happiness — like "99.9% of checkouts succeed within 2 seconds this month." You deliberately don't pick 100%, because users can't perceive the last sliver of reliability through their own flaky wifi, yet each extra nine costs roughly ten times more to deliver — 99.9% allows 43 minutes of failure a month, while 99.99% allows four, which is less time than a human needs to wake up and respond. The gap between your target and perfection is your error budget: 43 minutes of allowed badness that belongs to the team. While budget remains, you ship features and take risks freely; when it's spent, launches pause and everyone works on reliability until it recovers — a rule both sides agreed to before anyone was upset, so the number settles the argument instead of the loudest person. The budget even fixes alerting: instead of paging on every blip, you page only when the budget is burning fast enough to threaten the month. Reliability stops being a vibe and becomes a currency you measure, spend, and defend.
Your service has a 99.95% monthly SLO. This month it suffered: (a) a 5-minute full outage, (b) 40 minutes at 10% error rate, (c) 2 hours at 1% error rate. How much of the error budget is spent? Do the math.
Budget: 0.05% of 43,200 minutes = 21.6 minutes of full-outage-equivalent. Spend, weighting each incident by the fraction of requests failing: (a) 5 min × 100% = 5 min. (b) 40 × 0.10 = 4 min. (c) 120 × 0.01 = 1.2 min. Total: 10.2 of 21.6 minutes — about 47% of the budget, so you're still shipping. But notice the shape of the reasoning: partial outages convert to full-outage-equivalents by multiplying duration by error fraction, and a "small" 10% degradation for 40 minutes cost almost as much budget as a total 5-minute outage. If it's day 10, you've burned half a month's budget in a third of a month — worth a ticket and a look at what's leaking, before it becomes a freeze.
Take the WhatsApp-style chat system you scoped in p0c01's exercise. Instrument it: name three SLIs, commit to one SLO with the math, and say exactly what pages a human.
SLIs: (1) message delivery success rate — fraction of accepted messages delivered to online recipients within 5 seconds; (2) end-to-end delivery latency p99 for online-to-online sends; (3) connection success rate — fraction of clients that can establish and hold a session. SLO: 99.9% of accepted messages delivered within 5s over 30 days — that's a 43-minute-equivalent monthly budget, defensible by one on-call rotation without multi-region heroics. Pages: fast burn on the delivery SLI (budget going in hours = fire) and connection success collapsing (users locked out entirely). Explicitly not paging: CPU, memory, queue depth, one node down — those are dashboard diagnostics, and replication (p1c08) exists precisely so one dead node isn't user pain. If a dying node does hurt users, the delivery SLI pages within minutes anyway — symptoms catch what cause-lists miss.
Critique this paging list from a real on-call rotation: (1) CPU above 80% on any host, (2) disk above 70%, (3) any pod restart, (4) p99 latency of the checkout flow above 3s for 10 minutes, (5) nightly backup job failed. Which pages survive, and where does the rest go?
Only (4) survives as a page — it's a symptom, sustained, and actionable. (1) and (3) are causes with no inherent user pain: a healthy autoscaled system runs hot and restarts pods routinely; both become dashboard panels consulted during a real incident. (2) is a cause but a predictive one — demote to a ticket sized by urgency: "disk full in ~4 days" is a morning task, though "full in 2 hours" earns a page because it's about to become a symptom. (5) becomes a next-morning ticket — nothing the on-call does at 3am improves a failed backup, but letting it rot invites GitLab's famous 2017 incident, where the outage revealed every backup mechanism had been failing quietly for ages. General rule: pages for user pain now, tickets for trouble on a schedule, dashboards for everything else.
Your search page fans out to 200 shard servers and waits for all of them. Each shard answers within 20ms 99% of the time, but takes ~800ms 1% of the time. Estimate the fraction of searches that take ~800ms, and propose two mitigations.
Probability all 200 shards are fast: 0.99²⁰⁰ ≈ 0.13. So about 87% of searches run at roughly 800ms — the "rare" slow case is the common experience. Mitigations: (1) hedged requests — after waiting ~p95 (say 30ms), resend the straggler's query to a replica and take whichever answers first; this trims the tail dramatically for a couple of percent extra load. (2) Return best-effort partial results — answer with 198 of 200 shards at the deadline; for search, a 99%-complete result at 50ms beats a complete one at 800ms. (Also worth attacking the cause: why is 1% slow? GC pauses, cold caches, and a hot shard from p2c06 are the usual suspects.) The senior reflex on display: fan-out width times per-leg tail probability, before you build the page.
An engineer proposes a request-count metric labeled with: HTTP status, region, endpoint route template, user_id, and request_id — "so we can debug everything from one metric." You have 20M users. Which labels stay, which go, and where do the evicted ones live?
Stay: status (~10 values), region (~5), route template (~50 — the template like /users/{id}, never the raw path, which is unbounded). That's ~2,500 series: cheap forever. Go: user_id (20M values) and request_id (unbounded — one value per request, the worst possible label) — either one multiplies every existing series by its cardinality, so user_id alone takes 2,500 series to 50 billion and the monitoring cluster dies of memory long before the bill arrives. The evicted labels live where per-event detail belongs: request_id becomes the trace ID propagated through headers, user_id goes in structured log fields and trace attributes. Then debugging "user 4711 gets errors" is a log query, and "which region's error rate jumped" stays a millisecond metric lookup — each pillar doing its own job.
Second week on the creator-dashboard team at a video app. The dashboard shows every creator how many views their videos got. The number comes from a nightly job: at 2am, a batch job reads the whole day's view events from storage, counts them per video, and writes totals to a database. Simple. Reliable. Boring.
Then product asks for one small change: "Creators want to see views for the last hour, updating live."
Fine, you think — run the batch job more often. But the job takes three hours to chew a day of data, and even a trimmed-down version can't run every minute. The deeper problem isn't speed. Batch processing assumes a finish line: "here is yesterday's data, all of it, go." Views never stop arriving. You're being asked to compute an answer over data that never finishes.
So you try the obvious hack: on every view, increment a Redis counter. It works in the demo. Then reality arrives. A client retries a failed upload and the view counts twice. A phone that was offline uploads three hours of buffered views into the wrong hour. A counting bug ships on Tuesday, and on Friday someone asks you to recount Tuesday — recount from what? The counter only knows the present. Every patch you bolt on — dedup, late-event handling, some way to replay — is you rediscovering, badly, what stream processing systems already solved. This chapter hands you the real machinery.
Batch works on bounded data — a dataset with a beginning and an end. The contract is beautiful: the input is complete before you start, so you can sort it, join it, count it, and emit the answer. If the job crashes, you rerun it on the same input and get the same answer. Completeness is free; you just wait for the day to end.
A stream is unbounded — no end, ever. Any answer you give is provisional: "the count so far." And that exposes the trade that runs through this whole chapter: latency vs completeness. Wait longer before answering and your answer covers more of the data. Answer sooner and you risk having to revise it when stragglers show up. Batch sits at one extreme of this dial — wait until tomorrow, be complete. Streaming lets you pick any other point on the dial, and every mechanism you'll meet below (watermarks, lateness, windows) is a knob on exactly this dial.
| Batch | Stream | |
|---|---|---|
| Input | Bounded — complete before you start | Unbounded — never complete |
| Answer | Final | Provisional, may be revised |
| Latency | Hours (whenever the job runs) | Seconds |
| Failure recovery | Rerun the job | Restore state + replay (we'll get there) |
| Typical tools | Spark, BigQuery, a SQL script | Flink, Kafka Streams |
The first industry answer to "we need fresh numbers and correct numbers" was the lambda architecture (a name coined by Nathan Marz, creator of the Storm stream processor). Run two pipelines over the same events. A batch layer recomputes everything from raw history every few hours — slow but trustworthy. A speed layer processes events as they arrive — fast but approximate. A serving layer stitches the two together: batch results for everything up to this morning, speed-layer results for the gap since.
It works. And it has a cost that anyone who has run one will name through gritted teeth: every piece of business logic exists twice, in two codebases, on two frameworks, with two sets of bugs. Change how you count a "view"? Change it in the Spark job and the streaming job, deploy both, and pray they agree. When they don't — and they eventually don't — you get two dashboards with two different numbers and a meeting about which one is lying.
The kappa architecture is the deletion of that second codebase. The observation: the reason lambda needs a batch layer is recomputation — fixing mistakes by re-reading history. But if your events live in a replayable log (Kafka, from p1c11: an append-only file where every event has an offset and any consumer can rewind), you don't need a separate batch system to re-read history. You start a second copy of your streaming job at offset 0, let it chew through the past at full speed, write into a fresh table, and swap when it catches up. One codebase. Recomputation becomes an operational routine, not a parallel universe.
In 2025+ kappa is the default, and the justification is one sentence: stream processors got good enough that the batch layer stopped earning its second codebase. Modern engines give you durable state, exactly-once semantics, and replay at millions of events per second — the correctness properties lambda's batch layer existed to provide. You should still be able to say when the old answer wins, though. That's next.
Around 2010, LinkedIn's data infrastructure was a tangle: Hadoop jobs for batch analytics, a zoo of custom pipes shipping data between systems, and real-time needs (who viewed your profile, news feed) served by one-off systems. The team — led by Jay Kreps — built Kafka (open-sourced 2011) to replace the tangle with a single idea: every event goes into a durable, replayable, ordered log, and every system reads from it. In 2013 Kreps wrote the essay "The Log: What every software engineer should know about real-time data's unifying abstraction," arguing the log isn't plumbing — it's the backbone the whole data architecture should hang off. Then in 2014, having run lambda-style systems and eaten the maintain-everything-twice tax firsthand, he published "Questioning the Lambda Architecture," where he proposed the kappa alternative: keep enough history in the log, and recomputation is just your stream job replaying it. LinkedIn also built the Samza stream processor to consume those logs, and Kreps co-founded Confluent to commercialize Kafka. A decade later, "keep the events in a log, make everything a stream job over it" is the default shape of new data platforms — an architecture that started as one company deleting its duplicate code.
Kappa being the default doesn't make batch dead — it makes batch a choice you should defend with numbers. Batch wins when latency genuinely doesn't matter, because then it's strictly cheaper and simpler. A nightly billing rollup over 2 TB of events is one Spark job on spot instances for a few dollars, with no watermarks, no state management, no on-call stream job — and if it fails, you rerun it. Weekly ML training data, monthly invoices, compliance reports: nobody needs those in 200 milliseconds. The senior framing at the whiteboard: "This output is consumed once a day, so I'll batch it — streaming here buys latency nobody asked for, at the cost of a 24/7 stateful service someone gets paged for." Streaming for the dashboards, batch for the paperwork.
Strip the branding off Flink or Kafka Streams and here's what a stream processor is: a for-loop over an infinite list, with durable state. You could write the naive version today: a Kafka consumer that reads view events one by one and updates a hash map of video_id → count. Congratulations, that's a stream processor. Now watch it die in production. The process crashes: the map is gone, and where do you resume reading? Traffic grows past one machine: who owns which keys of the map? An event arrives late: which hour does it belong to?
Everything a stream processing framework does is making that for-loop crash-proof and shardable. The hash map becomes keyed state: the framework partitions the stream by key (same consistent-hashing idea as p1c09), so all events for video_42 land on the same worker, and that worker's slice of the map is stored durably, not just in memory. The "where do I resume" question becomes checkpointing, which gets its own section. Kafka is the substrate that makes all of it possible — because the log lets any reader rewind, the processor gets cheap time travel, and recovery becomes "rewind and replay" instead of "hope."
Quick sizing instinct, because numbers decide whether any of this is even hard: a video app does 50M views/day — about 600 events/sec average, call it 5,000/sec at peak. A single modest Flink or Kafka Streams job eats that without noticing. State: 10M videos times ~100 bytes of counter state is 1 GB — fits on one machine. So for this system, throughput is not the interesting problem; correctness is. Late data, duplicates, and recovery are where the design lives. (At 4 billion events/sec — hold that thought for the Alibaba story — the sizing math flips and throughput becomes the design.)
A user watches your video on a train. At 11:58 the train enters a tunnel; the phone loses signal and buffers the view event. At 12:40 signal returns and the event uploads. Question: which hour's count does that view belong to?
There are two clocks in every pipeline. Event time is when the thing actually happened: 11:58, stamped on the event by the phone. Processing time is when your server got around to seeing it: 12:40. A counter you increment on arrival counts by processing time — the view lands in the 12:00–13:00 bucket, which is simply wrong. For a rough ops graph, who cares. For billing, ad reporting, or anything a finance team reads, event time is the only defensible answer, and interviewers deliberately pick those domains.
But event time creates a nasty puzzle. If events can arrive 42 minutes late, when do you dare say "the 11:00–12:00 count is done"? Wait forever and you're a batch job again. This is the latency-vs-completeness dial, and the industry's knob for it is the watermark: a marker flowing through the stream that asserts "we believe all events with event time up to T have now arrived." When the watermark passes the end of a window, the window fires — computes its result and emits it. A common recipe: watermark = the latest event time seen so far, minus a slack for out-of-orderness. Slack of 30 seconds means you're betting events arrive at most 30 seconds out of order.
A watermark is a bet, and bets lose. Choose the slack from data, not vibes: measure the arrival-delay distribution. Say p99.9 of your events arrive within 30 seconds of their event time — but mobile has a long tail of minutes to hours (tunnels, airplane mode, dead batteries). You cannot hold a live dashboard hostage to the last phone in a tunnel, so you layer the defenses:
A teacher collecting homework after a school trip. The deadline is Friday (the window closes). She knows some kids' parents mail it in, so she waits until Monday before grading — that's the watermark: "by Monday, I believe everything that's coming has come." One envelope arrives Tuesday; she's kind, reopens the grade book, and adjusts — that's allowed lateness. An envelope arriving in July goes into a folder for the school office to sort out — that's the side output. Where the analogy breaks: the teacher grades once and amends once. A stream processor might fire and re-fire a window several times, and every downstream system has to be built to accept "the count for 11:00 changed."
"Count of views" over an endless stream is meaningless without a "per what." Windows are how you chop an unbounded stream into finite questions. Three shapes cover essentially every interview and most of production:
Every window above is state: partial counts sitting in the processor's keyed store, waiting for the watermark. Kill the process and that's potentially millions of half-finished windows gone. The recovery story is what separates a real stream processor from a counter living in one process's memory, and it has two halves.
Half one: checkpoints. Periodically — say every 30 seconds — the framework snapshots all state to durable storage (S3, HDFS), along with the log offsets that state corresponds to. The naive way is to pause the world, dump everything, resume: unacceptable at any real throughput. The elegant way is the Chandy-Lamport idea (1985, and Flink's checkpointing is a descendant of it): take a consistent snapshot of a running system without stopping it. The source injects a barrier — a special marker — into the stream, right after, say, offset 4,182. The barrier flows along with the data. When an operator sees the barrier, it snapshots its own state at that moment and forwards the barrier downstream. Since every operator snapshots at the same logical point — "I have processed everything before the barrier and nothing after" — the collected pieces form one consistent picture of the whole pipeline, taken while the pipeline kept running.
A restaurant manager wants an exact count of orders in progress without closing the kitchen. She slips a red card into the ticket rail between two orders. Each station — grill, fry, plating — keeps working, and the moment the red card reaches it, that station writes down its current tickets and passes the card on. No two stations pause, yet everything written down describes the same instant in the flow of orders: everything before the card, nothing after. That's a barrier snapshot. Where the analogy stretches: with multiple input streams, an operator must wait until the card arrives on all of them before writing (barrier alignment) — the grill station waiting for the card from both the dine-in rail and the delivery rail.
Half two: recovery = restore + replay. On a crash, the framework restores every operator's state from the last completed checkpoint and rewinds the log to that checkpoint's offsets — possible only because Kafka retains the events. The job then replays the gap between checkpoint and crash at full speed. Checkpoint every 30 seconds and your worst-case recovery replays 30 seconds of events; checkpoint every 10 minutes and a crash at minute nine replays nine minutes. That's the knob: checkpoint interval trades runtime overhead against recovery time — another decision a number should drive.
Notice what recovery just did: events between the checkpoint and the crash get processed twice — once before the crash, once on replay. Inside the framework that's fine; restoring state rewound the counts too, so internal state ends up right. The problem is the sink — the database or topic you already wrote results into before crashing. Replay will write those results again. Your dashboard double-counts.
Flink's end-to-end exactly-once closes this hole by making sinks transactional: results are written inside a transaction that only commits when the checkpoint completes, coordinated by a two-phase commit (2PC — the same protocol, with the same coordinator-stuck failure modes, you met in p2c01). Crash mid-interval and the uncommitted writes vanish with the failed checkpoint; downstream readers never see them. It genuinely works. And it has real costs: results become visible only when checkpoints commit (so end-to-end latency is roughly your checkpoint interval), the sink must support transactions, and you've invited 2PC's operational sharp edges into your pipeline.
Here's the pragmatic line, and I'll commit to it: at-least-once delivery plus an idempotent sink usually beats true exactly-once. Make the sink write an upsert keyed by (video_id, window_start) — "set the 11:00 count for video 42 to 1,847." Replay after a crash just overwrites the row with the same value. No transactions, no commit-latency coupling, no 2PC, and the observable result is exactly-once even though delivery wasn't. This is p2c02's lesson wearing streaming clothes: exactly-once processing is a property you engineer at the destination, not a gift the pipe hands you. Reach for transactional exactly-once only when the sink genuinely can't be idempotent — appending to a stream someone else consumes, or triggering side effects like payments, where "written twice, same value" isn't a concept that saves you.
Alibaba's Singles' Day (November 11) is the largest shopping event on Earth, and its famous real-time GMV dashboard — GMV is gross merchandise value, the running total of money spent, the giant number journalists watch tick upward all day — is a stream processing job. Around 2016, Alibaba forked Apache Flink into an internal version called Blink to harden it for search, recommendations, and that dashboard; in January 2019 they went all-in, acquiring data Artisans — the company Flink's creators founded, later renamed Ververica — for a reported €90 million, and contributing Blink's improvements back upstream. The scale numbers are the point: during Double 11 in 2020, Alibaba's Flink deployment peaked at about 4 billion events per second, roughly 7 TB of data per second, feeding dashboards executives and media were staring at in real time. At that scale, every mechanism in this chapter stops being theory: checkpointing must not stall the pipeline, watermarks must not hold windows hostage, and recovery must replay minutes of history at a rate that dwarfs most companies' total traffic. The same year they began running batch and stream workloads on the one engine — the kappa idea, "one codebase for both," executed at the most extreme scale on record.
Two weeks after launch you discover the sessionization logic has a bug: sessions split wrongly at midnight. Every session metric for 14 days is wrong. With a counter that only knows the present, this is a shrug and an apology. In the kappa world it's a runbook: deploy the fixed job as a second instance reading from the offset of 14 days ago, writing to a new table. It replays history at full speed — replay typically runs at 10–50x the live rate, so two weeks of data might take a few hours. When it catches up to now, flip readers to the new table, retire the old job. Same trick for schema evolution: want a new field aggregated? Replay history through the new code and the past gets the new field too, as if it had always existed.
This is the property you're really buying with the log-plus-stream-job architecture, and it comes with a bill you must name: your recompute horizon equals your log retention. Keep 7 days of Kafka and you can only replay 7 days; the bug found on day 14 is only half-fixable. That's why real platforms use tiered storage (Kafka offloading older segments to object storage) or archive every event to a data lake — cheap S3 bytes are what keep the "replay anything" promise honest. Retention is a budget line, not a footnote.
Calibration, because this topic has a wide staff-level ocean you should not drown in. This chapter is also the concept backbone for two Part 3 problems — Ad Click Aggregation (p3c13) and Metrics at Scale (p3c18) — where everything above gets assembled under interview pressure.
L5 must volunteer, unprompted:
Staff-only garnish (fine to mention, wasteful to derive): Chandy-Lamport's proof and unaligned-checkpoint variants, per-partition watermark generation and idle-source handling, RocksDB state-backend tuning, incremental checkpoint internals, the exact Kafka transaction protocol inside the 2PC sink. Name-drop at most; an L5 loop rewards the committed stance above far more than a lecture on any of these.
Stream processing rarely arrives as its own question. It ambushes you inside "design ad-click counting" or "design a metrics system," and the interviewer probes whether your numbers can be trusted. A passing senior answer names the two clocks, commits to a window and a late-data policy, and explains recovery without being asked. Probes you should expect verbatim:
Explain to a junior engineer, in five or six sentences, why "count views per hour" is hard on a stream — covering event time vs processing time and what a watermark is.
There are two clocks on every event: when it actually happened (event time) and when our servers received it (processing time), and they disagree whenever a phone is offline — a view from 11:58 can arrive at 12:40. If we count by arrival, that view lands in the wrong hour, so for anything involving money we must count by event time. But then we never know for sure when an hour is "complete," because a straggler could still be out there. A watermark is the system's moving bet: "I believe everything up to 12:05 has arrived now" — and when that bet passes the end of an hour, we publish that hour's count. Events that lose the bet and arrive later are handled by policy: update the published number, shunt them to a side channel for reconciliation, or drop them if the use case tolerates it. So the answer is never "the exact count instantly" — it's a trade between answering fast and answering completely, and the watermark is the dial.
Pick the window (tumbling / sliding / session) and justify in one sentence each: (a) revenue per calendar hour for finance, (b) "requests in the last 5 minutes" for an autoscaling trigger evaluated every 30 seconds, (c) average time users spend browsing per visit.
(a) Tumbling 1h — each event in exactly one bucket, cheapest, and finance thinks in calendar hours anyway. (b) Sliding, 5-minute size, 30-second step — a tumbling boundary could split a spike and delay the scale-up; note the cost: each event now lives in 10 windows. (c) Session with an inactivity gap (say 30 minutes) — "a visit" has no fixed length, so the data itself must define where the window ends.
Your watermark is "max event time seen minus 60 seconds," and the window is [12:00, 12:05). Events arrive in this order (event time → arrival time): A 12:01→12:01, B 12:04→12:04, C 12:03→12:05, D 12:05:30→12:05:45, E 12:02→12:08. When does the window fire, what count does it emit, and what happens to each event?
The watermark reaches 12:05 only when an event pushes max-event-time to 12:06 — D's 12:05:30 gets it to 12:04:30, so the window is still open when C arrives; A, B, C are all safely in. Whenever a later event lifts the max past 12:06, the window fires with count 3 (A, B, C — D belongs to the next window). E has event time 12:02, inside the window, but arrives after the watermark passed 12:05: E is late. With allowed lateness ≥ its delay, the window re-fires with count 4; otherwise E goes to your side output or is dropped, per policy. The lesson: the firing moment depends on the events you see, not the wall clock — an idle stream can hold a window open, which is exactly the idle-source problem real systems must handle.
Batch or stream? Commit and justify with one number each: (a) monthly invoice generation, (b) fraud checks on card payments, (c) the training dataset for a weekly-retrained recommendation model, (d) a live ops dashboard for the on-call engineer.
(a) Batch — consumed once a month; a stream job would run 720 hours to produce one artifact. (b) Stream — the decision is only worth anything within ~1 second, before the transaction completes; a nightly batch catches fraud 12 hours after the money left. (c) Batch — the consumer (weekly training) defines the freshness bar; anything fresher than weekly is wasted. (d) Stream with sliding windows — an on-call engineer needs to see a spike within seconds, not at the next batch run. Pattern: the consumer's read frequency, not the data's arrival rate, decides.
What breaks if you set the checkpoint interval to 10 minutes on a job processing 100K events/sec, and the job crashes 9 minutes after the last checkpoint? Walk the recovery, then say what an at-least-once + idempotent-upsert sink does to the user-visible damage.
Recovery restores state from the 9-minute-old checkpoint and rewinds the log: 9 min × 100K/sec = 54 million events to replay. At a 5x replay rate that's roughly 2 minutes of catch-up during which outputs are stale. All 54M events are processed twice; with a naive incrementing sink, counts double for that span. With idempotent upserts keyed by (key, window), replay overwrites rows with identical values — the user sees briefly stale, never wrong, numbers. If 2 minutes of staleness is too much, shorten the interval and pay checkpoint overhead more often: the interval is a recovery-time budget, and you should be able to say its value and why.
Your ads dashboard has one number at the top: unique viewers today. Not impressions — people. The PM wants it per campaign, live, refreshed every few seconds. You have 2 billion impression events a day across 4 million campaigns, and the obvious query — SELECT COUNT(DISTINCT user_id) — takes minutes on yesterday's data, never mind live.
Fine, you think. Keep a hash set of user IDs per campaign, in memory, and report its size. Do the math before you build it. A big campaign reaches 20 million distinct users. That's 160MB of raw 8-byte IDs — and a real hash set costs 3–5x raw, so call it 600MB. For one campaign. Your top thousand campaigns alone blow past half a terabyte of RAM, and there are 3,999,000 more behind them. Exactness has a price, and the price is proportional to the data.
Now the strange part. There is a data structure that answers "how many distinct users?" in 12 kilobytes — twelve, not a typo — with an error under 1%, no matter whether the true answer is ten thousand or ten billion. Your top thousand campaigns: 12MB total instead of 500GB. And when the counting is split across 8 shards, the 8 little structures merge into one correct answer with a single cheap operation.
The catch is in the fine print: the answer is approximately right, and you must know exactly which way it can be wrong. This chapter is about that trade — 99% right for 1% of the memory — and about a second, related trick: how a thousand machines agree on who's alive without any boss keeping a list. Both give up certainty, on purpose, in a direction they can defend.
Every exact structure you know — hash set, hash map, sorted index — remembers the items themselves. Memory grows with the data. That's fine until the data is a billion uniques, and then "have I seen this URL?" costs tens of gigabytes to answer honestly.
Probabilistic structures (people say sketches) remember a small, fixed-size fingerprint of the data instead. The fingerprint can't reconstruct the items, but it can answer one specific question about them — "seen before?", "how many distinct?", "how often?" — approximately. The senior discipline is that "approximately" is never hand-waved. Every sketch has a known error direction: which way it lies, how much, and never the other way. You pick a sketch the way you pick a tolerance on a machined part: because you computed the slack is acceptable here. Three sketches cover 95% of interviews, one question each.
Start with the pain. In the caching chapter you learned the miss path: not in cache → hit the database. Now imagine requests for keys that don't exist at all — misspelled URLs, deleted users, or an attacker spraying random IDs. Every one is a guaranteed cache miss (there's nothing to cache — the DB returns nothing), so every one lands on the database. This is called cache penetration, and it turns your cache into a decoration. What you want is a cheap gatekeeper that answers: "does this key even exist?" — for a billion keys, in RAM, in nanoseconds.
A Bloom filter is that gatekeeper. It's a long array of bits, all starting at 0, plus k hash functions (typically around 7). To add an item: hash it k ways, set those k bits to 1. To query: hash it the same k ways and look. If any of the k bits is 0 — the item was definitely never added, because adding it would have set that bit, and bits are never cleared. If all k bits are 1 — the item was maybe added. Maybe, because other items' bits can overlap yours by coincidence.
Memorize the asymmetry, because interviews probe it directly: false positives yes, false negatives never. "Not present" is a guarantee. "Present" is a probability. And note the corollary: you can't delete from a Bloom filter — clearing a bit might erase someone else's evidence. (A "counting Bloom filter" swaps bits for small counters to allow deletes, at several times the memory. Name it and move on.)
A Bloom filter is a club doorman with a strange memory: he never forgets a face he has actually seen, but sometimes a stranger "looks familiar." If he says "never seen you" — certainty, walk away. If he says "you look familiar" — probably a regular, occasionally a lookalike, so the club still checks the guest book (the database) to be sure. The analogy breaks in one spot: real familiarity fades with time. A Bloom filter's never does — it only accumulates, which is exactly why an over-full filter becomes the problem (more on that in the exercises).
The number to carry into the room: about 10 bits per item for a 1% false-positive rate (the exact figure is 9.6 bits with 7 hash functions; nobody wants the derivation). Halve the error every time you add ~5 more bits per item: 0.1% costs ~14.4 bits.
So: a billion URLs in a crawler's seen-set. Exact, as a hash set of even just 8-byte hashes: 8GB raw, 25–30GB as a real in-memory set. As a Bloom filter at 1% FP: 10 billion bits ≈ 1.2GB. That's the whole pitch: 20x less memory, and the 1% of errors all point in one direction you chose because you can afford it.
Akamai's CDN servers cache web objects on disk. Their researchers measured something embarrassing: roughly three-quarters of all requested URLs were fetched exactly once and never again — "one-hit wonders." Every one was written to disk cache, never read, then evicted to make room for the next never-read object. The fix, published in Maggs and Sitaraman's paper on Akamai's algorithms, was a Bloom filter as a bouncer for the cache itself: record every requested URL in the filter, but only write an object to disk on its second request — when the filter says "seen before." One-hit wonders never touch the disk. Disk writes fell by nearly half, and the hit rate went up, because space wasted on single-use objects now held things people actually re-requested. A false positive just means some object gets cached on its first hit — zero harm. Billions of URLs, a few gigabytes of RAM, no correctness risk: a sketch in exactly the right seat.
Second question, and the ads dashboard is where it bites: unique viewers per campaign, live, across 4 million campaigns. A Bloom filter can't help — it answers "seen before?", not "how many distinct?". Counting uniques exactly means remembering every ID you've seen, and that price is brutal — a campaign with 20 million distinct users runs ~600MB as a real hash set, so the top thousand campaigns alone pass half a terabyte.
HyperLogLog (HLL) counts distinct items in a fixed ~12KB with a typical error under 1%. The intuition is a bar bet. Hash every user ID — a good hash makes the output look like random coin flips. Now watch for rare patterns: a hash starting with one zero bit happens half the time; starting with twenty zero bits happens once in a million. So if the rarest pattern you've ever observed is 20 leading zeros, you've probably seen around a million distinct values. Duplicates change nothing — the same ID gives the same hash — which is precisely why HLL counts distinct for free.
One maximum is a fragile estimate (one freak hash wrecks it), so HLL runs thousands of these experiments at once: the first 14 bits of the hash pick one of 16,384 buckets, each bucket remembers the longest zero-run it has seen, and a clever average across buckets gives the estimate. 16,384 buckets × a few bits each ≈ 12KB, standard error 0.81%. That's the whole thing — in an interview you say "longest-run-of-zeros trick, thousands of buckets, averaged," and stop.
You walk into a room and ask: "everyone who flipped coins today — what's the longest heads streak anyone got?" If the best streak in the room is 5, it's a small room. If someone got 22 heads in a row, thousands of people must have been flipping — a streak that rare doesn't show up in small crowds. You estimated the crowd without counting anyone. Where the analogy breaks: one lucky liar ruins a single-streak estimate, so HLL runs 16,384 separate rooms and blends them — the blend is what buys the 1% error.
In practice you rarely build one: Redis ships HLL as a native type. PFADD campaign:123 user_9871 on every impression, PFCOUNT campaign:123 for the dashboard, and — the part that matters for distributed systems — PFMERGE to combine counters. BigQuery, Snowflake, Spark, and Druid all expose an equivalent under names like APPROX_COUNT_DISTINCT.
In April 2014, Salvatore Sanfilippo (antirez) added HyperLogLog to Redis 2.8.9 as a first-class data type. His design notes are a small masterclass in the trade: a dense HLL is exactly 12KB and counts to billions with 0.81% standard error, while small counters use a sparse encoding of a few dozen bytes, so a million mostly-small counters stay cheap. He named the commands PFADD, PFCOUNT, PFMERGE — the "PF" honors Philippe Flajolet, the French researcher who invented the algorithm and had died three years earlier. Reddit's engineering team later described rebuilding post view counts on exactly this: storing every viewer ID per post was off the table at their scale, so view events stream through Kafka into per-post HLLs in Redis, and the "N views" number on every Reddit post became a sub-1%-error estimate. Nobody can tell 1,024,590 views from 1,031,000 on a screen — but the memory bill can tell a set from a sketch.
Third question: not "seen?", not "how many distinct?", but "how often did X occur?" — hashtag frequencies, hot keys, top requested URLs. An exact hash map of counters grows with the number of distinct keys: fine for thousands, painful for hundreds of millions of live keys per hour.
A count-min sketch (CMS) is a small grid: d rows (say 4), each with its own hash function, and w counters per row (say 100,000). To record an event for key X: each row hashes X to one of its cells and increments it. To read X's count: look up its cell in each row and take the minimum. Why min? Because collisions can only inflate a cell — other keys land on it and add — never deflate it. Every row overestimates or is exact; the min is your least-contaminated view. So a CMS overestimates only, never underestimates. A 4×100K grid of 4-byte counters is 1.6MB — for any number of distinct keys, forever.
One limitation matters: a CMS can't tell you which keys are big — you bring a key, it returns a count. To get "top K hashtags" you pair it with a small min-heap: for every incoming key, ask the CMS for its estimate; if it beats the smallest of your current top K, swap it in. Sketch answers "how big?", heap remembers "who". That pair is the standard heavy-hitters design, and it's the spine of the Top-K chapter (p3c12) — here you just need to recognize it.
Here's how this topic actually gets graded. Anyone can say "use a Bloom filter." The L5 answer is three clauses in one breath: the structure, the direction it lies, and why that lie is affordable in this seat. "Bloom filter over existing keys — false positives cost one wasted DB read, false negatives can't happen, so no user ever gets a wrong 404. About 1.2GB for a billion keys at 1%."
And the rule cuts both ways. If the data fits exactly, use the exact structure. Ten thousand product SKUs needing counts? A hash map — under a megabyte, zero error, done. Reaching for a count-min sketch there is a flagged red flag: it says you memorized a tool, not a trade-off. The decision procedure is just arithmetic: size the exact answer first, out loud. If it fits in RAM on one box with headroom, exact wins. Sketches enter when the exact number stops fitting — and never for anything where wrong-by-1% means wrong-money (billing, ledgers, inventory: exact, always, as chapter p2c02 hammered home).
| Structure | Question it answers | Memory | Error direction | Merge across shards |
|---|---|---|---|---|
| Bloom filter | Have I seen X? | ~1.2GB per 1B items @ 1% FP | "Maybe yes" can be wrong; "no" never is | Bitwise OR |
| HyperLogLog | How many distinct? | ~12KB, any cardinality | ±~1%, either side | Per-bucket max |
| Count-min sketch | How often did X occur? | A few MB, any key count | Overestimates only | Cell-wise add |
| Plain hash map/set | Anything, exactly | Grows with data | None | Ship all the data |
Look at that last column — it's the quiet superpower. Take the ads dashboard again: impressions are sharded across 8 Kafka partitions and 8 consumers (chapter p1c09 made sure of that). Each consumer counts distinct users locally. Now produce one global number. Sum the 8 counts? Wrong — the same user appears on multiple shards, and you double-count. Exact fix? Every shard ships its full ID set to one machine: you're back to hundreds of gigabytes moving over the network.
Sketches merge losslessly. Two HLLs with the same bucket layout: take the per-bucket max, and the result is exactly the HLL one machine would have built seeing every event — merging costs 12KB of network, and the error doesn't compound. Blooms merge by OR-ing bits; CMS grids by adding cells. Merge is commutative and associative, so region rollups and hour-into-day rollups just work. One subtlety worth volunteering: replayed events (streams get replayed; chapter p2c09) are harmless to a Bloom or an HLL — same item, same bits, same buckets — but inflate a CMS, which counts occurrences. "HLL is replay-safe, CMS isn't" is a sentence that visibly separates levels.
Switch problems. Same trade, one level up: gossip gives up certainty about who is alive, on purpose, in a direction it can defend, and buys constant per-node cost with it. You run a 1,000-node Cassandra-style cluster. Every node needs to know which peers are up, which are down, and who owns which token ranges (the slices of the keyspace each node serves) — otherwise it can't route requests. Who keeps that list?
A central health-checker is the obvious answer and the wrong one at this scale: it's a single point of failure (checker dies → the cluster is blind), and "who health-checks the health-checker" never ends. All-to-all heartbeats — everyone pings everyone every second — is 1,000 × 999 ≈ a million messages per second of pure overhead, growing with N². A consensus store like ZooKeeper works for tens of nodes, but a thousand watchers hammering one quorum on every membership change becomes its own outage.
Gossip protocols do what rumors do. Every second, each node picks 1–3 random peers and exchanges what it knows — a compact membership list: node, status, and a version number per entry. Newer versions overwrite older. That's the whole protocol, and the math is the magic: information spreads like an epidemic, roughly doubling its audience every round, so it reaches all N nodes in about log₂(N) rounds. A thousand nodes: ~10 rounds, seconds. A million nodes: ~20 rounds. And each node's cost is constant — a few small messages per second — no matter how big the cluster gets.
Office gossip needs no all-hands meeting. One person hears the news at the coffee machine, mentions it to two colleagues, each mentions it to two more — by lunch, the whole floor knows, and nobody was "in charge" of announcing it. Where the analogy breaks, instructively: office rumors mutate as they spread. Gossip protocols prevent mutation by versioning every fact — each entry carries a counter, and everyone keeps only the highest version they've heard. The rumor can't distort; it can only get fresher.
One boundary to volunteer before the interviewer asks: gossip gives you eventually consistent facts — perfect for membership, load hints, schema versions, anything that tolerates a few seconds of disagreement. It does not give you agreement. Electing a leader or committing a transaction still needs real consensus (chapter p2c03). "Membership via gossip, coordination via Raft" — one sentence that shows you know which tool holds which weight.
Gossip spreads small, fresh facts. But replicas also drift in data: a node was down for an hour, missed writes, and the hinted handoffs (chapter p1c08) that were supposed to catch it up got dropped. Now two replicas of the same partition silently disagree. Comparing them row by row means shipping the whole dataset over the network.
A Merkle tree makes the comparison cheap. Each replica builds a tree of hashes over its key ranges: leaves hash small ranges of rows, parents hash their children, up to one root hash covering everything. The replicas exchange roots. Equal → the terabytes match, proven in one round trip. Different → recurse only into the children that differ, down to the handful of ranges that actually diverge — then stream just those rows. Logarithmically many hash comparisons instead of the data itself; this is exactly how Cassandra's repair and Dynamo's anti-entropy work. The honest cost, raised unprompted: building the tree reads and hashes the data on disk, so repair is an expensive scheduled operation, not a free background hum — ops teams run it off-peak and stagger it across nodes.
Membership has a second half: detecting failures. Naive version: A pings B; no answer → A declares B dead, gossips it, and the cluster starts rebalancing B's data. Count the ways that goes wrong: B was in a garbage-collection pause. The network path between A and B hiccupped. A's own machine was overloaded and never sent the ping. Each false alarm moves gigabytes of data — a false positive here isn't a wasted lookup, it's a self-inflicted outage.
SWIM — the protocol under HashiCorp's Consul and Serf (via their memberlist library, with Lifeguard extensions) — layers three defenses:
One knob remains: how long is the timeout? A fixed number is wrong twice. Too tight, and every GC pause triggers a false alarm; too loose, and real failures take ages to detect. The phi-accrual failure detector (used by Cassandra and Akka) replaces the binary timeout with a dial: it records the history of heartbeat arrival gaps per peer and outputs a continuous suspicion score — "given how this node usually behaves, how improbable is this silence?" A node that always heartbeats like clockwork gets suspected after a short quiet spell; a node on a jittery link earns more patience. The detector adapts per node, automatically, which is exactly what a hand-tuned timeout can't do.
In 2007 Amazon published the Dynamo paper — the shopping-cart store built to survive anything, and the origin of half of Part 1's vocabulary: consistent hashing, quorums, hinted handoff, Merkle-tree anti-entropy, gossip membership. One of its authors, Avinash Lakshman, then joined Facebook, where inbox search needed a write-heavy store no single failure could stop. With Prashant Malik he built Cassandra: Dynamo's distribution layer under Bigtable's table-shaped data model, open-sourced in 2008. Look inside a Cassandra cluster today and you find this entire chapter in production: every node gossips its state once per second with up to three random peers; liveness comes from a phi-accrual detector, not fixed timeouts; replica drift is repaired by comparing Merkle trees; and every SSTable carries a Bloom filter so point reads skip files that definitely lack the key. No master node anywhere — which is why the design survived from Facebook's inbox to Netflix and Apple, who run some of the largest Cassandra fleets in the world today. One database, four ideas from this chapter, all still doing their original jobs.
Calibration matters, because this chapter's material splits cleanly into "must be fluent" and "must merely recognize" — and candidates routinely invert it, deriving HLL math while never mentioning what a false positive costs.
Working level — you volunteer this, unprompted: Choosing sketch vs exact from the numbers: size the exact answer first, and if it fits, say a hash map wins. The three-clause pattern: structure, error direction, cost of an error in this seat. The sizes: ~10 bits/item at 1% for Bloom, 12KB for HLL, megabytes for CMS. The merge story when the design is sharded — that's the moment sketches go from trivia to architecture. And the standard seats: penetration guard, SSTable skip check, crawler seen-set, unique counters, heavy hitters.
Recognition level — one good sentence and move on: "Membership and failure detection via SWIM-style gossip — probe, indirect probe, suspicion with incarnation numbers — like Consul's memberlist; replica repair via Merkle-tree anti-entropy like Cassandra; detection timeouts adaptive via phi-accrual." That sentence, delivered casually while designing something else, is a full senior signal. Nobody at L5 is asked to reproduce SWIM's message formats.
Staff-only garnish — skip unless asked: deriving the 9.6-bits formula, HLL's harmonic mean and small-range corrections, CMS width/depth error bounds, Lifeguard's self-awareness tweaks, phi's log-scale definition. Knowing these exists; reciting them costs you deep-dive minutes that score higher elsewhere.
This topic almost never arrives as its own question — it surfaces inside crawlers, feeds, analytics, and any sharded design. A passing senior answer names the structure, its error direction, and its size in one breath, then wires the error direction to the seat. Probes you should expect:
Explain to a junior engineer, in four or five sentences, what a Bloom filter is, what it can and cannot promise, and one place you'd actually use it.
A Bloom filter is a big array of bits that answers "have I seen this item before?" using a fraction of the memory of a real set — about 1.2GB for a billion items instead of tens of gigabytes. Adding an item sets a few bits chosen by hash functions; checking an item looks at those same bits. If any bit is 0, the item was definitely never added — that's a hard guarantee, because bits are never cleared. If all bits are 1, the answer is only "probably yes," because other items' bits can overlap — around 1% of the time it says yes to something it never saw. So you use it where a false "yes" is cheap and a false "no" would be a bug: for example, in front of a database to instantly reject lookups for keys that don't exist, where a rare false positive just means one wasted database read.
Sizing drill: your crawler will see 10 billion URLs. Size the seen-set as (a) an exact store and (b) a Bloom filter at 0.1% false positives. Then state — in one sentence each — what a false positive and a false negative would mean for the crawler.
(a) Even storing only an 8-byte hash per URL: 80GB raw, realistically 250GB+ as a live hash set — a disk-backed store, not RAM. (b) 0.1% FP needs ~14.4 bits per item → 144 billion bits ≈ 18GB — fits in RAM on one large box. A false positive says "seen" for a brand-new URL, so you skip ~1 in 1,000 new pages (tunable; if that matters, confirm "maybe" answers against the exact disk store). A false negative would mean re-crawling duplicates forever — and it cannot happen. That error direction is why the Bloom filter is the front line and not the other way around.
Pick the structure, and justify with error direction: (a) "has this email ever registered?" checked before an expensive DB lookup; (b) daily active uniques, aggregated across 200 app servers; (c) the 10 most-requested API endpoints this hour, out of millions of distinct endpoints-with-parameters; (d) live counts for 20,000 feature flags.
(a) Bloom filter of registered emails — a false positive costs one DB read; a false "not registered" would be a correctness bug and can't happen. (b) HLL per server, PFMERGEd centrally — per-bucket max gives exactly the union's sketch; summing 200 distinct-counts would double-count users who touched multiple servers. (c) Count-min sketch + min-heap — CMS bounds memory against millions of keys and only ever over-counts; the heap remembers who's on top. (d) Plain hash map — 20,000 counters is under a megabyte; any sketch here is the memorized-tool red flag. Numbers first, structure second.
What breaks: you sized a Bloom filter for 1 billion items at 1% FP, but the crawl kept going and it has now absorbed 4 billion. What happens to its behavior, and what's the fix?
The bit array saturates — at 4x design load most bits are 1, and the false-positive rate doesn't degrade gently, it climbs toward "says yes to everything." For a crawler that means silently skipping a huge fraction of genuinely new URLs: the system doesn't crash, it quietly stops discovering — the worst kind of failure. Fixes: rebuild bigger (blooms can't resize in place — bits don't map over), run generational filters (a fresh filter per time window, query recent ones, retire the oldest), or chain progressively larger filters (scalable Bloom). The senior move is naming the monitoring: track the fill ratio (fraction of set bits) and alert before saturation — this failure is invisible in every other metric.
Gossip judgment: in a 300-node cluster, one node with long GC pauses gets declared dead every few minutes, each time triggering a rebalance that moves gigabytes. Name the three mechanisms from this chapter that stop the flapping, and what each contributes.
(1) Indirect probes (SWIM): other nodes try their own network paths to the pausing node, filtering out cases where the prober, not the target, had the problem. (2) Suspicion with incarnation numbers: a missed probe makes the node suspect, not dead; when the pause ends, the node gossips "alive" with a higher incarnation number, overriding the suspicion cluster-wide before any rebalance starts. (3) Phi-accrual detection: the detector learns this node's heartbeat history — pause pattern included — and raises its patience for this peer specifically. Together: false alarms become suspicion windows that quietly expire, and the gigabyte-moving machinery engages only for real deaths. (Also worth saying: fix the GC pause.)
Two years ago someone at your company sharded the users table four ways — user_id % 4, clean and simple, Chapter 1.9 stuff. It worked. The company grew. Now shard 2 is at 85% disk, its p99 is twice everyone else's, and capacity planning says you have four months before it falls over. The fix is obvious: go from 4 shards to 16. The problem is everything else: hundreds of millions of rows have to physically move to different machines, the routing logic in every service has to change, and at this exact moment about 40 million people are logged in.
The old answer was a maintenance window. Put up a banner — "down for maintenance Sunday 2–6 a.m." — stop the writes, run the scripts, check the counts, bring it back up. Half the internet was built this way.
So pick your window. 2 a.m. where? You have users in forty countries; 2 a.m. in Virginia is mid-morning in Mumbai and evening in Sydney. There is no hour when your traffic sleeps, because your traffic is the planet. And even if there were: Chapter 2.8 gave you an SLO of 99.95%, which is a budget of about 22 minutes of downtime per month. A four-hour window spends a year of error budget on one schema change — and this is the first of three migrations on this quarter's roadmap.
This is also a real interview question, almost word for word. At the senior level it arrives as: "You need to change the shard key on a live table with zero downtime. Go." No hints, no menu. There is a standard professional answer to this — a pattern so established that Stripe, GitHub, and Shopify have each written it down in public. By the end of this chapter you'll be able to walk it end to end, with a rollback plan at every step.
Every zero-downtime migration — new column, new table, new database, new shard layout, even a whole new datastore — is the same four-phase movie. Learn it once and you can improvise any variation.
Two rules make this pattern safe, and both are worth saying out loud in an interview. First: every phase is independently deployable and independently revertible. You never ship two phases in one deploy, because then a problem in either forces you to unwind both. Second: you decide how to get back before you go forward. Amazon's vocabulary is useful here: a two-way door is a decision you can walk back through; a one-way door isn't. The whole art of zero-downtime migration is converting one giant one-way door (the big-bang cutover) into a corridor of small two-way doors.
It's moving to a new house while living your normal life — no hotel, no vacation days. First you get the keys to the new place and set up mail forwarding, so everything new arrives at both addresses (expand: dual writes). Then, over a few weekends, you move the old boxes room by room (migrate: backfill). Before giving up the old lease you walk both houses with a checklist — is everything really there? (verify: shadow reads). Only then do you change your official address, and you still keep the old keys for a month in case you find an empty drawer that shouldn't be empty (contract: flip, watch, then let go). Where the analogy breaks: your furniture doesn't change while you're moving it. Your data does — thousands of times per second. That's the entire technical difficulty, and it's what the next three sections are about.
Phase one ships the new destination — new column, new table, the 16 new shards — plus code that writes every change to both old and new. Sounds like two lines of code. It is not, because the second write can fail, and the two writes are not one atomic action.
The discipline that keeps you safe has three parts:
Even with discipline, application-level dual writes have a known flaw: the app can crash between the two writes, and now the stores disagree and nothing recorded it. If you want to close that gap properly, you already know the tool from Chapter 2.1: the outbox pattern — commit the change and an event describing it in one transaction, and let a relay apply the event to the new store. Or go one level lower with CDC, change data capture (Chapter 2.9): tail the database's own write-ahead log or binlog — the append-only journal where the database records every row change — and replay those changes into the new store. CDC can't miss a write, because it reads the same journal the database itself trusts. The senior framing: "I'll dual-write from the app for speed of shipping, or from CDC for correctness — for a money table I'd insist on CDC or outbox, because a silent gap between two non-atomic writes is not a risk I'll carry on a ledger."
Dual writes only cover changes from now on. Everything written before — years of history — has to be copied by a batch job, while the system runs at full traffic. Do the arithmetic before anything else, because the number changes the design: 2 billion rows at a polite 10,000 rows per second is 200,000 seconds — a bit over two days. At that scale the backfill is not a script someone babysits; it's a small production service with three non-negotiable properties:
updated_at before writing, or your backfill will happily overwrite tonight's data with last year's.That last bullet deserves a slow-motion replay, because it's the most common way this pattern goes wrong in practice. The backfill reads row 42 at 14:00:00. At 14:00:01, a user updates row 42 — the dual-write puts the fresh value in both stores. At 14:00:02 the backfill, holding its stale 14:00:00 copy, writes row 42 into the new store. Blind upsert: the new store now has old data, and no error was raised anywhere. This is why "idempotent" alone isn't enough — the write must be conditional: apply only if what I hold is newer than what's there.
The backfill finished. Row counts match. Are the stores identical? You don't know. Counts don't catch a million rows with a mangled timezone. Before you point users at the new store, you make the system prove equivalence — continuously, on real traffic.
The tool is the shadow read (also called a dark read): on some fraction of real read requests, the app reads both stores, serves the old store's answer to the user — the user experience cannot regress, because the old answer is by definition today's behavior — and compares the two results in the background. Every disagreement increments a mismatch metric and logs a sample. Now you have a dashboard that measures the one thing you care about: do the two stores agree on the reads production actually performs, which weights hot rows exactly as much as production does. Pair it with an offline checker — chunked checksums over both stores — to cover the cold rows nobody read this week.
The gate is strict and worth stating as policy: cut over only when mismatches are approximately zero and have stayed there for days. Not "low." A steady 0.01% mismatch rate on a billion daily reads is a hundred thousand wrong answers a day, and a steady rate is never random noise — it's a systematic bug (a null handled differently, a timezone conversion, an encoding) sitting in a specific slice of your data. Find the class, fix it, re-backfill that slice, watch the metric fall to zero. Then move.
In 2017 Stripe published "Online migrations at scale," the post that made this pattern a named industry standard. They needed to restructure hundreds of millions of Subscription objects — customers were moving from one-subscription-each to many — on data that processes real money, where a wrong row isn't a stale feed item, it's a wrong charge. Their recipe is exactly this chapter's spine, stated as four phases: dual-write to the old and new representations; then move all read paths over; then move all write paths; then remove the old data. Each phase shipped separately and was validated before the next began, with the backfill running against a live production collection. Their framing is worth quoting in an interview because it names the payoff precisely: each step is small enough to reason about, and the system answers traffic normally the entire time. The migration took months of calendar time — and users never saw any of it. Months of patient, boring, reversible steps is what "zero downtime" costs, and at Stripe's stakes that trade is obviously right.
Cutover day is deliberately boring, because the switch is not a deploy — it's a feature flag: a runtime setting your code checks on each request, flippable in seconds without shipping code (you met flags in Chapter 2.8's rollout discussion). You flip reads to the new store for 1% of traffic. Watch p99 latency, error rate, and the mismatch metric — which you keep running, now comparing in the other direction. Then 10%, then 50%, then 100%, watching at each step. If anything looks wrong at 1%, you've exposed 1% of users for a few minutes, and the rollback is one flag flip. Compare that to discovering the same bug at 100% after a big-bang cutover with no way back.
One ordering rule does a lot of quiet work here: flip reads before you flip writes. While both stores are still being written, reads can move back and forth freely — both copies stay current, so the read flag is a perfect two-way door. Once reads are 100% on the new store and stable for days, the new store becomes the source of truth, and then you change the write path. Even here, the careful move keeps a door open: reverse the dual write — new store first and authoritative, old store still updated best-effort — so that for a few more weeks even a late-discovered disaster has a path back to the old store without data loss. Then, and only then: remove the flag, stop writing the old store, snapshot it, and drop it.
Now the trap. Every step of contract is optional in the sense that the system works without it — and that's exactly why teams stop halfway. Reads are flipped, everything's green, the sprint ends, priorities shift. The dual writes stay on. Quietly, forever.
Six months later you have two sources of truth: every write costs double, every schema change must be made twice, the copies drift (the repair queue stopped being watched in March), and during every incident someone asks "which store is right?" — and gets two answers. Naming this failure mode unprompted is a genuinely strong signal — it's the difference between knowing the pattern and having lived with one: "Dual-write is a state I pass through, not a state I operate in. Contract gets a calendar date before expand ships, and the migration isn't 'done' on the tracker until the old path is deleted."
| Phase | What ships | Who serves reads | How you get back |
|---|---|---|---|
| Expand | New store + dual writes (or CDC/outbox relay) | Old | Flag off dual writes; delete new store; nothing happened |
| Migrate | Backfill job: chunked, rate-limited, checkpointed, idempotent | Old | Pause or abandon the job; old store never depended on it |
| Verify | Shadow reads + mismatch metric + offline checksums | Old | Nothing to revert — you simply don't proceed until ~0 |
| Contract | Read flag 1%→100%; then write flip (reversed dual-write); then delete old | New | Flip the flag back in seconds; after write flip, reversed dual-write keeps old current |
Not every migration is a new datastore. The most common migration in any company's life is humbler: add a column, change a type, add an index — on a MySQL table with 400 million rows. And the naive move, a plain ALTER TABLE, is a small production incident you schedule yourself. Historically MySQL rebuilt the whole table while blocking writes — hours of read-only on a big table. Modern "online DDL" — schema changes the database applies while still serving reads and writes — is better but still bites: the ALTER can wedge behind one long-running query on a metadata lock, and then every query behind it queues — a traffic jam where nobody honks, the app just stops. And on replicas the ALTER applies as one giant serial operation, so your read replicas fall an hour behind while it runs.
The fix is the expand/migrate/contract pattern in miniature, automated by a tool. First generation: Percona's pt-online-schema-change creates an empty copy of the table with the new schema (a "ghost" table), backfills it in chunks, and keeps it current using triggers — little stored procedures on the original table that mirror every insert, update, and delete into the ghost. It works, but the triggers run inside your production write transactions: every write to the real table now pays extra latency and lock contention, you can't pause a trigger, and on an already-hot table that overhead arrives exactly when you can least afford it.
GitHub's schema migrations were frequent — their workflow ran continuous changes to a MySQL fleet serving the busiest developer site on earth — and trigger-based tools kept hurting on their hottest tables: the trigger overhead landed inside production write transactions, migrations couldn't truly pause, and aborting one was messy. So in 2016 their database team built and open-sourced gh-ost, whose announcement leads with the design goal: a triggerless online schema migration tool. Instead of triggers, gh-ost subscribes to MySQL's binary log — typically from a replica, so even the listening cost leaves the primary alone — and replays row changes onto the ghost table while copying old rows in small chunks. Because gh-ost owns the entire flow, it can do what triggers never could: throttle itself when replica lag climbs, pause completely on command, resume, and hold before the final table-swap until a human says go. The cut-over itself is an atomic rename — the app sees the old table one instant and the new one the next. It's the chapter's whole philosophy compressed into one tool: never block the live path, copy behind it, stay controllable, keep the exit open.
Postgres users have gentler defaults — adding a nullable column is instant, and indexes can build concurrently — but the same discipline applies to anything that rewrites a big table. And one rule is universal, worth volunteering in any interview that touches schemas: never couple a schema change to a code deploy. Expand the schema first (new column, nullable, ignored by old code), deploy code that handles both shapes, migrate the data, and only later contract by dropping the old column. If the deploy rolls back, the schema still fits the old code. Sound familiar? It's the same four phases, applied to one table.
The version interviewers reach for when they want senior-or-above signal is a live users table sharded four ways that has to become sixteen. Changing the shard key or shard count is the hardest migration there is: the data's home address changes, every row moves, and the routing layer — the code deciding which shard holds user 7,204,911 — must agree with reality at every instant. Here's the full walk, with the rollback at each step, exactly as you'd narrate it at a whiteboard.
Step 0 — exploit the math. Going 4 → 16 is a multiply-by-4, and that's not an accident. With modulo routing, every user on shard user_id % 4 = 1 can only land on new shards 1, 5, 9, or 13 — x % 4 = 1 forces x % 16 into that set. Each old shard splits cleanly into four children; no row ever crosses between old shards. (This is the same instinct as consistent hashing in Chapter 1.9: choose a scheme where growth means splitting, not reshuffling.) Say this in the interview — it cuts the problem's size by 75% before any machinery appears.
Expand. Stand up the 16 new shards, empty. Teach the router both layouts: old (mod 4) stays authoritative for reads and primary writes; a flag turns on dual writes to the new layout. For resharding at scale, prefer CDC over app-level dual writes — tail each old shard's binlog and route each change to its new home. One relay per old shard, no app code in the write path, no crash-between-two-writes gap. Rollback: flag off, truncate the new shards. Cost so far: hardware.
Migrate. Four backfill workers, one per old shard, each streaming its rows into that shard's four children — chunked, throttled on replica lag, checkpointed, idempotent, version-checked against the CDC stream's newer writes. At 10,000 rows/second per worker, the four streams together move 2 billion rows in about 14 hours — the clean four-way split is also a four-way speedup. Rollback: pause; the old shards never stopped being the truth.
Verify. Shadow reads on a sample of production traffic: route the read both ways, serve old, compare, count. Offline checksums sweep the long tail. Gate: ~0 mismatches, held for days. Rollback: still nothing to roll back — you just don't proceed.
Contract. Flip reads by flag: 1% of users routed to the new layout, then 10, 50, 100, watching p99 and mismatches at each stop — and if the graphs so much as twitch, flip back and diagnose at leisure, because the old shards are still fully current. After days of green, make the new layout authoritative for writes, keep a reverse CDC stream feeding the old shards for two more weeks as the final escape hatch, then decommission: snapshot the old shards, drop them, delete the mod-4 code path. Put the date for that last step on a calendar during expand — you know why by now.
A hospital moving to a new building doesn't pick a Sunday to wheel every patient across town and hope. It opens the new building empty, staffs it, and runs both in parallel. New admissions get charts in both systems (dual writes). Old records move ward by ward, checked against the originals (backfill + verify). Patients transfer a ward at a time, stable ones first (percentage cutover), and the old building keeps power, staff, and beds until the last transfer has been stable for weeks (the old store outlives the flip). The one thing a hospital never does is close the old building the same day the new one opens — and neither do you.
Shopify runs on sharded MySQL, with each merchant's shop living on one shard — and shards fill up unevenly, because one merchant going viral inflates a shard (Chapter 2.6's problem wearing a different shirt). So they built Ghostferry, an open-source Go tool that is the migrate phase productized: batch-copy a shop's rows to the target shard while tailing the binlog to replay concurrent changes, with verifiers that check data integrity before, during, and after the move. On top of it, moving a shop between shards is a routine operation: the copy runs in the background at full traffic, and the move ends with a brief write-lock on that one shop — seconds, not hours — while the final binlog entries drain and the router repoints. That's the mature endgame: rebalancing so routine it's just another background operation, and "downtime" scoped to one tenant for seconds. Now the dark twin. In October 2010, Foursquare ran user check-ins on two MongoDB shards of 66GB RAM each. Growth skewed; one shard blew past its RAM and started thrashing to disk. They tried to fix it live, under peak load — adding a third shard and migrating chunks off the hot one — but freeing 5% of the data didn't free 5% of the memory: the deleted rows left the remaining data fragmented across the same number of pages. The site went down for about eleven hours while they compacted the database offline, and MongoDB's CTO published the postmortem. Same operation as Shopify's — moving data off a hot shard — but attempted as an emergency instead of rehearsed as a routine. The lesson seniors quote: reshard while it's a project, because done as a rescue it inherits the emergency.
Calibration, so you neither undershoot nor gold-plate. At the senior bar, an interviewer hearing "we need to change this live" expects you to volunteer, unprompted:
Staff-level garnish — welcome for a sentence, but never at the cost of the list above: gh-ost's cut-over locking choreography, cross-region cutover sequencing on top of Chapter 2.5, building a generic migration platform with pluggable verifiers, formal reconciliation for financial data. The grading distinction is blunt: the senior list is how you execute one migration safely; the staff garnish is how you build the machinery so fifty teams can. Nail the first completely before touching the second — a candidate who name-drops Ghostferry but can't say why backfill needs a version check has it exactly backwards.
This topic is rarely its own interview; it's an ambush inside another one. You've designed something, it works, and then: "the users table needs a new shard key — zero downtime. Go." A passing senior answer walks expand → migrate → verify → contract concretely for this system, with a rollback per phase and numbers on the backfill. Probes to expect:
A junior asks: "We need to change a column's type on a 500M-row production table. Why can't we just run ALTER TABLE tonight?" Explain the zero-downtime way in five or six sentences.
A plain ALTER on a table that size can lock or lag the database for hours, and we have users on it around the clock — there is no "tonight." So instead of changing the table in place, we build the change next door: add the new column, and ship code that writes every update to both the old and new column, while reads still use the old one. A background job then copies the 500M historical rows across in small, resumable batches, careful never to overwrite a row the live writes already updated. Then we prove the two columns agree: on real traffic we read both, serve the old one, and count every mismatch — we only proceed when that counter sits at zero for days. The switch itself is a feature flag: point 1% of reads at the new column, watch, ramp to 100%, and if anything looks wrong we flip back in seconds. Weeks later, once nothing has moved, we delete the old column — and only then is the migration actually done.
Your service has a 99.9% monthly SLO. A teammate proposes a 3-hour maintenance window for a storage migration, arguing "99.9 isn't that strict." Do the arithmetic and give a verdict — plus the one scenario where you'd approve a window anyway.
99.9% of a 30-day month leaves about 43 minutes of allowed downtime. Three hours is ~180 minutes — a bit over four months of error budget spent on one planned event, before a single unplanned incident this quarter. Verdict: no window; use expand → migrate → contract, which costs calendar time but zero availability. The exception is judgment, not dogma: an internal tool with tolerant users and no SLO, or a system with genuinely idle hours (a single-country B2B app at 3 a.m. local), where a short window is the cheapest correct answer. Seniors scale the ceremony to the stakes — in both directions.
An engineer implements dual writes as: write the NEW store first ("it's the future, it should be primary"), then the old store; if either fails, fail the request. Name two concrete problems this ordering and policy create during the migration window.
First: failing the user's request when the new store errors couples your production availability to an unproven shadow system — the warming-up copy can now take down the site, which inverts the whole point of keeping the old store authoritative. Second: writing new-first means a crash between the writes leaves data in the new store that the old (source-of-truth) store never saw; since all reads serve the old store, the user's successful-looking write is invisible to them — and worse, the backfill or verifier may later "confirm" the new store's phantom row. Correct discipline: old store first and synchronous (its failure fails the request, same as today), new store second and best-effort with the failure logged for repair — or skip app-level ordering entirely and drive the new store from CDC/outbox so there's nothing to race.
You're backfilling 2B rows and can sustain 8,000 rows/sec without pushing replica lag past your threshold. How long does the backfill take? Halfway through, the job crashes and, separately, the verifier later finds ~50K rows where the new store holds older data than the old store. Diagnose both.
2,000,000,000 ÷ 8,000 = 250,000 seconds ≈ 2.9 days — so this was always going to be a multi-day service, not a script. The crash is the expected case, not the surprise: with per-batch checkpoints ("done through key X") the job resumes from X and loses minutes; without them it restarts from zero and, if its writes are blind upserts, re-clobbers days of dual-written data as it re-runs. The 50K stale rows are the classic race: the backfill read those rows, a live dual-write updated them, then the backfill's stale copy landed last. The fix is a conditional write — apply only if the incoming version/updated_at is newer than what the new store holds — and the remedy now is to re-copy exactly those rows with the condition in place, then let the verifier confirm the count falls to zero.
Sketch the full plan (phases, verification, rollback per step) for changing a 4-shard user table to 16 shards with modulo routing — then answer: which users' rows can end up on new shard 9, and why does that matter operationally?
Plan: (1) Expand — 16 empty shards; router learns both layouts; CDC relays from each old shard's binlog route live changes to the new homes; rollback = stop relays, truncate. (2) Migrate — one backfill worker per old shard streaming into its four children: chunked, lag-throttled, checkpointed, version-checked upserts; rollback = pause. (3) Verify — shadow reads on sampled traffic (serve old, compare, count) plus offline checksums; gate at ~0 mismatches for days. (4) Contract — read flag 1→10→50→100%, then writes flip with a reverse relay feeding the old shards for two weeks, then snapshot and drop, on a date scheduled back in step 1. Shard 9: only users with user_id % 4 = 1, because x % 16 = 9 forces x % 4 = 1 — new shard 9 is one of old shard 1's four children {1, 5, 9, 13}. Operationally this means every row's source is known and single: workers never contend, verification can run shard-pair by shard-pair, and a problem found on new shard 9 implicates exactly one backfill stream, not sixteen.
"Let's start with something simple. Design a URL shortener — like bit.ly." The interviewer smiles. You almost relax. Every course opens with this problem; you've seen a hundred diagrams for it. Client, server, database, done. What could possibly go wrong?
Here's what the smile means. The interviewer knows this problem is small — that's why they picked it. No celebrity fanout, no consensus, no petabytes. Nowhere to hide: every point comes from judgment — the numbers you run, the choices you commit to, the traps you spot in features that look trivial. Mid-level candidates recite the memorized diagram and stall when asked "why?" three times in a row. Senior candidates treat it as what it is: a small problem that deserves a senior answer, with pacing so clean the interviewer never has to steer.
This first walkthrough does double duty: it solves the shortener at the E5 bar, and it narrates the clock — where each phase lands in a 45-minute round — so you can watch chapter 0.1's framework actually run. One sentence per phase notes the time; that habit is half of what "drives the room" means.
Even a small problem gets scoped out loud: "I'll build: shorten a URL, redirect a short code to its long URL, optional custom aliases like bit.ly/my-launch, optional expiry, and lightweight click stats — counts, not dashboards. Out of scope: user accounts beyond an owner ID, paid analytics. Fair?"
Then the non-functional list, with numbers committed rather than fished for:
abc123 exists, an attacker learns nothing about abc124. Why this matters becomes a war story later.Clock check: five minutes gone, scope agreed, flag already planted on the read path — that's where the deep dives will go.
The whole product category exists because of a unicycle forum. In 2002, Kevin Gilbertson, a unicyclist from Blaine, Minnesota, got tired of long URLs wrapping and breaking in emails to his unicycling newsgroup, so he built TinyURL — the first notable shortener, and one of the oldest still running. It stayed hobby-sized until Twitter's 140-character limit made every long URL painfully expensive; in May 2009 Twitter made bit.ly (launched 2008) its default shortener, replacing TinyURL, and traffic exploded — competitor tr.im announced shutdown within months of losing that slot. Then Twitter built its own wrapper, t.co, and took the default back. That should have killed bit.ly. Instead the company had realized the redirect log is the product — every click says who clicked what, when, from where — and it survived by selling that analytics layer to marketers. Remember this at 301-versus-302 time: the boring response code is a revenue decision.
Run the numbers — they're about to overturn the assumption most candidates walk in with.
Now say the conclusion out loud: "This is a read-heavy, small-data system. Storage will never be the problem — 3 TB fits on one box — so I won't shard for data size; that would solve an imaginary problem. The real constraints are read QPS and latency, so the design is a cache in front of a key-value store, and everything else serves that." The estimate just made three decisions: no sharding for size, cache-first architecture, deep dives on the read path. Estimation doing its job — driving decisions, not decorating the whiteboard.
A URL shortener is a coat check. You hand over something long and unwieldy; you get back a tiny numbered ticket. Nobody thinks the hard part is closet space — a thousand coats fit in one room. The hard parts are the ticket numbers (two customers must never get the same one, even with three attendants working at once) and the pickup rush, a hundred times busier than drop-off. Closet space = storage: a non-problem. Tickets and the pickup queue = ID generation and the read path: the whole game. One break in the analogy — a coat check hands out tickets in order, and we specifically must not.
Two minutes, not ten — this section proves you can turn requirements into shapes, then gets out of the way.
POST /links with {long_url, custom_alias?, expires_at?} → 201 {code: "xK9fQ2a"}. Authenticated; accepts an idempotency key so a retried request doesn't mint two codes (chapter 2.2 habit — one sentence, earns a nod).GET /{code} → 302 with Location: long_url, or 404 unknown, or 410 Gone expired. Unauthenticated — the hot path.GET /links/{code}/stats → {clicks: 18234}. Owner-only.One table carries the system — links: code (primary key), long_url, owner_id, created_at, expires_at. Click counts live elsewhere, never written synchronously by the redirect — that's deep dive three. Notice the access pattern: every hot-path query is "the row for exactly this code." Pure key-value, no joins, no ranges. Hold that thought — it decides the database.
Draw the complete loop before going deep anywhere — both paths, traceable end to end.
Write path: client posts a long URL → Create API takes the next ID from its locally held block (leased a million at a time from a counter table — deep dive one), permutes and base62-encodes it, writes code → long_url, returns the short link. One conditional write, no retry loop.
Read path: GET /xK9fQ2a → redirect service checks Redis → on hit (target 95%+) responds 302 in a few milliseconds → on miss, reads DynamoDB, fills the cache, responds. Either way it fires a click event at a queue and forgets it.
Database: I'll commit to DynamoDB. The access pattern is a pure key-value get by primary key, which is exactly what Dynamo is; conditional writes give atomic "insert if not exists," which the alias race needs; built-in TTL handles expiry cleanup; and it scales reads without me operating replicas and failovers for what is, at heart, one GET endpoint. The cost: lock-in and no rich queries — which this system never needs. The honest alternative: Postgres plus the same Redis cache is defensible at 12K reads/sec; I'm choosing the smallest 3 a.m. surface. Clock check: minute twenty, the loop is closed, and I announce my deep-dive agenda instead of waiting for one: "The interesting problems are code generation, the redirect path, and click counting. Code generation first."
This is why the question exists. You must generate billions of codes that are (a) unique — a collision silently sends someone to the wrong page, (b) short — 7 characters, not a 32-character hash, (c) mintable by many servers at once without queueing on a central "next number please" service, and (d) non-sequential — no guessable series. Every naive strategy fails at least one of the four, and the interviewer knows which one yours fails. Spend your best ten minutes here.
First, the alphabet. Base62 — digits 0–9, a–z, A–Z — because all 62 characters are URL-safe without escaping; base64's + and / need percent-encoding, defeating the point of a short link. Seven characters gives 62^7 ≈ 3.5 trillion codes; our five-year forecast of 6 billion links uses 0.17% of that. Seven is enough forever — say so and move on.
Now, how do you hand those numbers out to a fleet? Walk the options like a candidate who rejected each for a reason:
What's inside the scrambler? Even "multiply by a large odd constant, take the remainder mod 62^7" is a perfect shuffle of the space. A Feistel network — a small reversible mixer borrowed from cryptography — is stronger, because its pattern is much harder to reverse-engineer from examples. Name it, don't derive it; the interviewer wants the property (bijection = shuffle = no collisions ever), not the algebra.
Leased ranges are raffle-ticket books. The organizer (counter table) doesn't hand out tickets one by one from a booth with a queue around the block — she hands each volunteer a book of 1,000 pre-numbered tickets and goes back to her chai. Volunteers tear off tickets at full speed: no line, no radio calls, no two people holding ticket 4711, because the books never overlap. The permutation is the twist a raffle lacks: every number passes through the same secret shuffle before printing, so no buyer can compute a neighbor's number from their own stub. Honest limit: a volunteer who loses a half-used book loses those numbers forever — true for a crashed server's stranded block too. We don't care: burning a thousand million-ID blocks wastes 0.03% of a 3.5-trillion space. Wasted IDs are free, and knowing they're free is the senior read.
Self-identified weakness, before the interviewer finds it: the permutation is obscurity, not cryptography — if the constant leaks, codes become enumerable again. So I never treat unguessability as a security boundary (a short code is a public name, never an access control — the abuse section makes this concrete), and I watch the 404 rate for anyone probing the space. Second weakness: the counter row is a single point for leases — but with day-level lease frequency and hours of runway per server, it can be down a long time before anyone notices. I'll take it.
The redirect response code looks like trivia. It's a business decision, and interviewers love it because the trade-off has no free lunch.
A 301 Moved Permanently tells the browser "this mapping is forever — remember it." Browsers cache it aggressively: the second click goes straight to the destination without touching your servers. Great for the bill; fatal for two requirements. Cached-away repeat clicks are invisible to counting, and a "permanent" answer can't expire or be taken down — a browser that cached the 301 keeps sending users to a link you've since disabled for malware. A 302 Found says "here's the destination today — ask again next time," so every click flows through you.
I'll commit to 302 (with explicit no-cache headers), because we committed to click stats and expiry, and both die under 301 caching. The honest cost: I volunteer to serve every repeat click myself — the whole 12K/sec — instead of letting the world's browsers absorb them. If the requirements flipped to no-analytics, immutable links, 301 becomes right and cuts real traffic; that's the version TinyURL could live with, and why bit.ly, whose post-Twitter business was the click log, needs every click to come home. The response code follows the requirements, not fashion.
Having volunteered for every click, serve them from memory. Cache-aside in Redis — the app asks the cache first and fills it itself on a miss, so the cache is an accelerator the system can live without: key = code, value = long URL plus expiry, TTL a day. Clicks follow a power law — a small fraction of links (the tweet going viral right now) draws most traffic — exactly the shape caches love: a modest cache holds the working set and 95%+ of redirects never touch the database. 12K/sec peak fits one Redis node with a replica; I'd still run two shards for failure isolation, not throughput. If one link goes truly nuclear — a hot key in the chapter 2.6 sense — a tiny in-process LRU for the top hundred links caps the damage. One more trick that pays for its sentence: negative caching. Cache "this code does not exist" for a few minutes too, so a scanner hammering random codes — or a typo'd link in a big group chat — burns cache, not database.
The trap: "increment a counter on every redirect" quietly turns your read-heavy system write-heavy — 12,000 counter writes per second so a dashboard nobody is watching stays current to the millisecond. Say the quiet part: click counts are decorative data. They tolerate delay and small loss. The redirect tolerates neither.
So the redirect service fires a fire-and-forget event — {code, timestamp} — onto a queue (chapter 1.11's machinery) and returns the 302 without waiting. A small consumer group aggregates per-minute counts and flushes batched increments — "code xK9fQ2a: +1,417" — every few seconds. Twelve thousand events per second becomes a few dozen database writes per second; dashboard counts run seconds behind reality, declared acceptable when we scoped "stats lite." If the product later demands uniques or top-K trending, that's the count-min sketch and friends from chapter 2.10 bolted onto this same event stream — the stream is the extension point, the real reason to build one at all.
The failure policy is the part that reads senior: if the queue is down, drop the events and keep redirecting. Emit a metric so we know we're undercounting, and move on. Backpressuring the redirect path to protect analytics would invert the system's priorities. Stats are the passenger; redirects are the plane.
Custom aliases. "I want short.ly/my-launch" is a uniqueness claim, and two users can claim the same word in the same millisecond. The classic bug is check-then-insert: read "is it free?", see yes, insert — and between your read and your write, someone else did the same dance. The fix is to never check at all: one atomic conditional insert — in DynamoDB, a put with attribute_not_exists(code); in Postgres, an insert against the unique primary key — and let the database referee. Exactly one claimant succeeds; the loser gets "taken, pick another." Same lesson as ticket booking in chapter 3.8 at one-thousandth the drama: uniqueness is enforced by an atomic write, never by a read followed by hope. Aliases share the namespace with generated codes, so add one structural rule — generated codes are exactly 7 characters, aliases must be 8+ — and the two can never collide by construction.
Expiry. Two mechanisms, and you need both — that's the interview point. Lazy expiry: the redirect path checks expires_at on every hit and answers 410 Gone when past-due — the correctness mechanism, working even if cleanup never runs. Cleanup: DynamoDB's built-in TTL deletes expired rows for free, but it's explicitly best-effort and can lag a day or two — precisely why the lazy check must exist. Lazy check for correctness, TTL for hygiene; neither alone is complete. (Also evict the cache entry on the 410 path, or Redis serves the dead link until its own TTL runs out.)
A URL shortener is a machine that makes any link look harmless. That's the product — and the attack. Phishers were among the earliest heavy users of shorteners, because short.ly/xK9fQ2a hides the sketchy domain it points to. Raise this unprompted: check submitted URLs against a threat feed like Google Safe Browsing at creation, re-scan asynchronously (attackers submit a clean page, then swap its content once shortened), and make takedown cheap — which our 302 choice enables and a cached 301 would sabotage. Rate-limit creation per user and per IP (chapter 1.12's token buckets), because a spammer's first move is minting a million links.
In 2016, researchers Martin Georgiev and Vitaly Shmatikov published a study with a perfect title: "Gone in Six Characters: Short URLs Considered Harmful for Cloud Services." The observation was brutally simple: token spaces this short — five to seven characters — are small enough to scan by brute force. They enumerated short links minted by goo.gl and OneDrive's 1drv.ms — and found "shared privately" documents behind short links were effectively public. Scanning surfaced live OneDrive documents, and from a discovered link they could often traverse to the rest of the account; about 7% of exposed accounts were writable by anyone — a ready-made mass-malware vector, since cloud folders sync down to their owners' machines. Scanned Google Maps short links revealed shared driving directions — enough, for many users, to infer home addresses. The fallout: Microsoft dropped URL shortening from OneDrive; Google moved Maps to much longer tokens. Two lessons, say both: entropy is a speed bump, not a lock — private things must authenticate at the destination, never hide behind a short code — and a spiking 404 rate is the sound of someone scanning your keyspace, which is why it belongs on the dashboard.
Close the way you'd run the service. One pass, unprompted:
Probes come fast on an opener, aimed at judgment density. Nearly verbatim:
| Decision | Chose | Over | The cost I accepted |
|---|---|---|---|
| Code generation | Leased ID ranges + bijective permutation | KGS pool; hash+truncate; raw counter | Permutation is obscurity, not security; crashed servers strand partial blocks (0.03% of space — free) |
| Redirect code | 302 + no-cache | 301 permanent | We serve every repeat click ourselves — analytics and revocability are paid for in QPS |
| Database | DynamoDB | Postgres + replicas | Lock-in, per-request pricing, no ad-hoc queries — fine, the access pattern is one key-value get |
| Click counts | Async queue + batched aggregation | Synchronous increment | Counts lag seconds and can undercount during queue outages — declared acceptable up front |
| Expiry | Lazy check + DynamoDB TTL sweeper | Either alone | Two mechanisms to reason about; TTL lags up to ~48h (correctness never depends on it) |
| Alias claims | Atomic conditional insert | Check-then-insert | None — this one is simply correct; check-then-insert is simply a race |
A junior asks: "A URL shortener is just a hash map with a website in front — why do interviewers keep asking about it?" Answer in five or six sentences, naming the one genuinely hard part.
You're right that the data is tiny — a billion links fit on one drive — and that's the first lesson: notice it's small and refuse to shard imaginary petabytes. What's not trivial is minting the codes: many servers must generate billions of unique 7-character codes without queueing at a central counter and without a guessable sequence, because guessable codes let strangers walk your keyspace. The trick: lease each server a big block of numbers so coordination is rare, then push every number through a reversible shuffle so output looks random but can never collide. Second lesson: the redirect is read-heavy and latency-critical, so a cache serves nearly everything, and click counting goes through a queue so it can never slow a redirect. Even the response code is a decision: 301 lets browsers cache the redirect and you lose analytics; 302 means every click comes to you and can be counted, expired, or revoked. Small system — but every piece is a judgment call, which is what the interview grades.
Requirement flip: click counts are now exact — billed to advertisers at ₹2 per click. Which parts of the design must change, and which survive?
The fire-and-forget click path dies: "drop events when the queue is down" is now dropping money. Billing needs durable accounting — persist the click event (acked queue write, or an outbox row) before returning the 302, deduplicate on an event ID downstream, reconcile daily (chapter 2.2's ledger thinking). The redirect gets slower and gains a dependency — the honest price, and saying so is the senior framing. Everything else survives untouched.
Your CEO read a blog post and demands 301 redirects "to cut serving costs." Write the two-sentence reply that commits, and name the product change that would make the CEO right.
Something like: "301 lets browsers cache the redirect, so repeat clicks never reach us — never counted, and a link we take down for malware keeps working in every browser that cached it; analytics and abuse takedown both break. If we ever drop click stats and promise immutable links, 301 becomes right and I'll make the switch that day." Commit, price it, name the flip condition — that last clause separates a preference from a judgment.
Now add multi-region: users in India, Europe, and the US should all see <50 ms redirects. Sketch the minimal extension. What's the consistency catch, and why is it almost harmless here?
Redirects are reads, so replicate the read path: Redis plus an async DynamoDB global-table replica per region, DNS/anycast routing users to the nearest. Creation stays single-region-writer — 120/sec needs no geo-distributed writes. The catch: a just-minted link takes a second or two to reach remote regions, so a Berlin user clicking a brand-new Mumbai link could 404. Almost harmless: creators read via their home region, cross-region clicks within seconds of creation are rare, and a 404 on a fresh code can fall back to the writer region. Bonus: pre-partition the ID space per region so even leases never cross an ocean.
What breaks first if link creation goes 1000x — say, an API partner starts minting 40,000 links/sec programmatically? Walk the write path and find the choke points in order.
In order: (1) The Safe Browsing check — an external synchronous call per create — chokes first; go async (create optimistically, scan within seconds, kill bad links after the fact). (2) Leases jump to one per ~25 seconds fleet-wide — still trivial, but worth 10M blocks. (3) Dynamo write throughput needs provisioning — a bill, not a design change; note the permutation accidentally guarantees no hot partition, since consecutive IDs scatter. (4) The real break: 40K/sec ≈ 100B links/month, so five years is 6 trillion links — overflowing 62^7's 3.5T space. The redesign is 8-character codes: the constant you defended at minute 22 is what 1000x actually breaks.
Estimation drill, 90 seconds: how much RAM does Redis need for a 95% hit rate, given the power-law click pattern? State your assumptions.
Assume today's traffic concentrates on recent and viral links: say 95% of the day's ~330M redirects hit the ~20M most active links — well under 1% of all 6B (power laws are steeper than intuition suggests). Entry: 7-byte code + ~200-byte URL + overhead ≈ 300 bytes. 20M × 300B = 6 GB — one modest Redis node. The meta-answer: cache sizing is a footnote here, and knowing when a sub-problem is a footnote is the skill being drilled.
The interviewer writes one line on the board: "Design search autocomplete — the suggestions under the search box as you type." Then the trap, delivered casually: "Assume Google-ish scale."
Most candidates hear "search suggestions" and think small. It's a dropdown. Ten strings. How hard can ten strings be? Then the first honest piece of math drops the floor out: the search engine answers one query per search, but autocomplete answers a query per keystroke. Whatever traffic the search engine handles, this little dropdown handles ten times more — and it has to be faster than search, because it's racing the user's fingers. A suggestion that arrives 300 milliseconds after the keystroke arrives after the next keystroke. It's useless.
This problem is a favorite at Meta and Google because it looks cute and is secretly brutal. It has a clean crux, it punishes "just use a database," and it rewards letting numbers make the decisions. Let's play it at the senior bar: commit, justify, find the fire early.
First move in the interview, out loud: shrink the problem — top completions for a prefix, at search-engine scale — and get agreement.
Functional:
Deliberately out of scope, and I'll say why: per-user personalization (your own history ranked first). Personalization makes every response user-specific, which destroys shared caching — and shared caching is going to carry half this system. I'll design the global system and show at the end where a "personalization lite" layer bolts on client-side. That's a product trade-off stated as an architecture decision, made out loud.
Non-functional, with numbers I'm committing to:
That last line is a CAP-flavored commitment (Part 1) made in product terms, and it decides several designs below.
One multiplication drives this entire chapter, so let's do it slowly.
5 billion searches/day ÷ 86,400 seconds ≈ ~60,000 searches/sec average; call peak 2×, ~120K/sec. Now the amplifier: the average query is ~20 characters, and the client asks for suggestions as the user types. Even after client-side tricks we'll add later (debouncing, caching), figure roughly 10 suggestion requests leave the device per search typed.
10 × 60K = 600,000 requests/sec average. Past a million at peak.
Sit with that. Autocomplete serves an order of magnitude more traffic than the search engine it decorates, with a tighter latency budget — and the budget is mostly pre-spent: on a mobile network the round trip alone can eat 50–80 ms of our 100, before the browser has drawn a pixel. What's left for the server is maybe 10–20 ms, a million times per second.
Autocomplete at Google didn't arrive as a roadmap item. In 2004 a Google engineer, Kevin Gibbs, built it in 20% time — the one-day-a-week personal-project budget Google handed engineers — and it shipped that year as an experiment on Google Labs, the company's public sandbox for half-finished ideas. Anyone could try it. And then it sat there. It took until 2008 for suggestions to become standard under the search box on google.com. Four years, for a feature you could demo in an afternoon. The delay wasn't product indecision: a Labs toy serving curious visitors and a system eating roughly ten times google.com's query traffic at lower latency than search itself are not the same machine. The idea was cheap; making it survive the scale was the entire job. That gap — trivial feature, brutal scale — is exactly what this interview question measures.
Imagine a shop assistant who must run over and offer help after every letter a customer writes on their shopping list, not every item. One customer buying five things generates fifty interruptions. That's autocomplete's life: the "real" work happens once, but the assistant sprints on every pen stroke. Any design where each sprint is expensive — a scan, a disk read — collapses under the sheer count of sprints. The only survivable design is one where the assistant already holds the answer in hand.
Storage estimate, because it gates where data lives. Keep the top ~2 billion distinct queries worth suggesting (the junk long tail adds nothing), and precompute an answer for every prefix of every kept query. Prefixes overlap heavily ("car", "care", "career" share "car"), so call it ~10 billion distinct prefixes. Each entry — ~20-byte key, 10 completions with scores, overhead — ~500 bytes. That's ~5 TB. Two conclusions: it doesn't fit on one machine, so we shard; and 5 TB across a fleet is comfortably RAM-sized, so nothing on the hot path ever touches disk. About 50 shards of ~100 GB, times 3 replicas — ~150 nodes. Big, but boring-big. The math just designed the storage tier.
Every candidate's first instinct is a table of queries with counts and:
SELECT query FROM searches
WHERE query LIKE 'ca%'
ORDER BY freq DESC LIMIT 10;
A B-tree index on query handles the LIKE 'ca%' part fine — it's a range scan (Part 1, indexes). The killer is the ORDER BY freq. The index gives rows sorted alphabetically; you want the top 10 by frequency. So the database must fetch every row in the range — millions for a two-letter prefix — sort them, and throw away all but 10. Per keystroke.
And here's the cruel correlation: the shortest prefixes are both the most expensive (widest ranges) and the most frequent — you can't reach "cats" without passing "c" and "ca". Your worst query is your hottest query, a million times a second. No index tuning escapes this; the scan-then-rank work is real. The fix isn't a faster database. The fix is to stop doing the work at request time.
Say the senior sentence out loud: "The read path can't afford to compute anything. So I'll precompute the answer to every possible question."
GET /suggest?q=cat
→ 200 {"q":"cat","s":["cat videos","cat food","cat memes", ...]}
Notes worth saying, then dropping: the prefix is normalized (lowercased, trimmed, Unicode-normalized) so "Cat" and "cat " hit the same cache entry; the response is a few hundred bytes; and the endpoint is read-only and cacheable, which we will exploit hard. The core data model is one logical structure:
prefix (string, normalized) → top-K completions [ (text, score) × 10 ]
That's it. The entire serving system is a giant read-only dictionary from prefix to pre-ranked list. Everything interesting is in how we build, shard, and refresh it.
Two loops, fully closed. The read path: keystroke → client cache and debounce → edge/CDN → suggest service → in-memory prefix shards → back. The write path: every executed search is logged → Kafka → two consumers: a batch job that periodically rebuilds the whole prefix→top-K world, and a streaming job that catches what's trending right now. The serving fleet merges both at read time.
Walk the read path like a request: a keystroke fires. The client checks its tiny session cache and applies a debounce. If a request goes out, it hits the CDN edge — and because everyone on Earth types through the same short prefixes, the edge answers a huge fraction of traffic a few milliseconds away, with a short TTL (~60 seconds) so trending stays fresh-ish. Misses reach the suggest service: hash the prefix, call the owning shard, get the precomputed list in under a millisecond, merge any trending overlay (below), return. No step computes anything meaningful. That's the whole point.
The write path gets its own deep dive — it's where ranking and freshness live. First, the crux.
This is the fire in this house: p99 under 100 ms, at a million requests per second, ranking billions of historical queries — which means the ranking must already be done before the request arrives. Everything else in this design is plumbing around one idea: precompute the top-K answers for every prefix and store the system as a read-only dictionary in RAM. Candidates who try to make request-time search fast (better indexes, faster databases, clever tries traversed per request with live re-ranking) are fighting physics. Candidates who move all the work to build time and make the read path a single hash lookup have found the crux. The interview question is secretly: do you know when to precompute?
The classic mental model is a trie — a tree where each node is one more character of prefix: root → c → ca → cat. The subtree below any node contains every query starting with that prefix, and to avoid scanning subtrees per request you cache the top-K completions at every node during the build. Lookups become: walk down, read the stored list. A beautiful teaching model — I'll keep it in my head.
But I won't deploy it, and here's my reasoning. Our data is 5 TB — it must be sharded. How do you shard a tree? By subtree: one machine owns everything under "a", another under "b"... and you've recreated Part 2's hot-partition problem — "s" and "c" subtrees are enormous and scorching, "x" and "z" idle, and rebalancing means moving subtrees. Trie lookups also chase pointers node to node, and the structure invites clever in-place updates that turn a simple system into a concurrent data-structure project.
My commitment: a flat hash table — prefix → top-K list — sharded by hash of the prefix. Because:
The cost I'm accepting: storing every prefix explicitly duplicates data a trie would share, roughly doubling memory. RAM is the cheapest thing in this design; I'll pay it. (A lever if memory hurts: cap stored prefixes at ~10 characters and let the client filter a longer stored list — traffic past 10 characters is a trickle.)
Self-identified weakness: hot prefixes. Hashing spreads distinct keys evenly, but not traffic — everyone typing anything starting with "c" hammers whichever shard owns key "c", tens of thousands of requests per second on one key. Two-layer fix, committed: short prefixes (length ≤ 2) are a rounding error in size — 36² ≈ 1,300 two-character keys, a few megabytes — so every suggest-service node keeps a full local copy and never calls a shard for them. For hot longer prefixes, run extra replicas of hot shards (hot-key replication, straight from Part 2's celebrity-problem playbook). And the CDN already absorbed most short-prefix traffic before it got here. Defense in depth against a completely predictable skew.
Where do the top-K lists come from? From the users: every executed search is a vote. The write path turns billions of votes into ranked lists.
Run it like a bakery. Overnight, the big ovens bake the day's stock of proven bestsellers — the batch rebuild, thorough and cheap per loaf, done while nobody's watching. But if a cake goes viral on TV at noon, you don't wait for tomorrow's bake: a small counter oven handles today's rush — the streaming layer. Customers see one shelf holding output from both ovens. The counter oven's cakes are rougher (approximate counts, short history) — fine, they only exist until tonight's big bake incorporates them properly.
Oven 1 — the batch layer. Search logs flow into Kafka; a daily MapReduce-style job aggregates weeks of history: count query frequencies with exponential recency decay (halve a query's weight every ~week, so dead memes sink without needing deletion), apply quality filters (drop queries below a minimum count — killing both junk and privacy-sensitive rare strings — plus misspellings, safe-search hits, and blocklist matches), then emit each prefix's top 10 by score, packed into immutable snapshot files per shard. Filtering at build time is deliberate: a blocklist applied once per build costs nothing; the same check per request would run a million times a second on identical data. Never do per-request what you can do per-build.
Oven 2 — the streaming layer. A daily build means breaking news wouldn't appear until tomorrow — failing the trending requirement. So a stream job (Part 2) consumes the same Kafka topic, counting queries in ~5-minute windows — sketch-based counting is fine here (a count-min sketch is a fixed-size table of counters that trades exact counts for tiny memory), because trending only needs "this is exploding," not an audited number. Queries whose current rate massively exceeds their historical rate get promoted into a small trending overlay: a few thousand hot queries and their prefixes. It's tiny, so we don't shard it — we push the whole overlay to every serving node every few minutes. At read time a node merges: score = batch_score + boost × trending_count, re-sorts ten-ish items, returns. Nanoseconds of work.
If you did Part 2's pipelines chapter, you've spotted the shape: a textbook lambda architecture — batch layer plus speed layer, merged at read. We argued there that kappa (one streaming path, replay to recompute) is usually the cleaner default. Why lambda here? Because the two layers genuinely compute different things: ranking wants weeks of decayed, heavily filtered history — natural as a batch job — while trending wants a five-minute window. When the speed layer fits in a pocket, lambda's usual sin — two full codebases that must agree — mostly disappears. Committed, with the reason attached.
Update cadence, committed: full rebuild daily, trending overlay every ~5 minutes, CDN TTL 60 seconds. Rebuild hourly? That's 24× the compute to freshen data the trending layer already covers. Make trending 10-second-fresh? The CDN TTL and browser cache would still smear staleness over a minute — real money spent improving a number the user can't see. The cadence lands where marginal freshness stops being visible.
Backend-minded candidates skip this, and it's a mistake — especially in Meta's product-architecture track, where "the client is a tier of your system" is house religion. Three client behaviors change our backend math:
Debounce. A fast typist emits a keystroke every ~80 ms. Firing per raw keystroke wastes half your traffic on prefixes already typed past; wait ~50–100 ms of typing silence first. This is where "10 requests per search, not 20" came from — the client already halved the fleet. Tune carefully: debounce too long and you've spent the latency budget standing still.
Cancel stale responses. You send "ca", then "cat". Nothing guarantees arrival order — the "ca" response can land after the "cat" response, and if the client renders whatever arrived last, suggestions flicker backward while the user types forward. The fix is small and mandatory: abort in-flight requests on new keystrokes, and tag each response with its prefix so the client drops anything not matching the current input. Volunteering this out-of-order bug — a distributed-systems race living inside a text box — is a cheap, genuine senior signal.
Cache and filter locally. Backspacing "cats" to "cat"? The answer for "cat" is in session memory — no request. And a client holding results for "cat" can filter them to serve "cats" instantly while the network catches up.
In 2010 Facebook's engineers published "The Life of a Typeahead Query," walking through how friend-and-page search-as-you-type actually worked — and the striking part is how much of the architecture lives before their servers. The moment a user merely focuses the search box, before typing one character, the browser prefetches that user's friends, pages, groups, and events into a local cache, so the first keystrokes resolve entirely on-device — zero network round trips. Only when local results run thin does a request go out, carrying the list of results already displayed so the backend won't resend them. Server-side it hits an aggregator: a stateless root fanning out to specialized leaf index services in parallel, merging results. Their typeahead is personalized by definition (your friends, not the world's top queries) — which is exactly why per-user data got pushed to the one cache that's per-user for free: the user's own browser. Put each piece of data in the cheapest tier that can own it.
Design something you'd be on call for. What breaks, unprompted:
Typeahead's follow-ups are well-worn. Expect these, near-verbatim:
| Decision | Chosen because | The cost you accepted |
|---|---|---|
| Precompute top-K per prefix (no request-time search) | Read path must be O(1) at 1M QPS with ~10 ms of budget | Freshness bounded by build/overlay cadence; storage duplicated across prefixes |
| Flat hash table over trie | Uniform hash-sharding; one-hop lookups; immutable offline rebuilds | ~2× memory vs a shared-structure trie |
| Lambda-shaped pipeline (daily batch + tiny streaming overlay) | The layers compute different things; the overlay is small enough to replicate everywhere | Two codepaths; scores briefly inconsistent as overlays roll out |
| Daily rebuild + 5-min overlay + 60 s CDN TTL | Matches the freshness users can perceive through the cache layers | Non-trending ranking shifts land up to a day late |
| Availability over freshness (serve stale, degrade to empty) | Autocomplete is an enhancement; latency is the product | Outdated or missing suggestions during incidents — accepted silently |
| Filtering at build time + tiny read-time suppression list | Per-build checks are ~free; per-request checks ×1M QPS are not | Non-emergency filter fixes wait for the next build |
| No server-side personalization | Per-user responses would zero out CDN and shared-cache hit rates | Personally irrelevant suggestions rank high; mitigated client-side only |
A junior asks: "Why does autocomplete need all this? Why not just query the database for matching strings when the user types?" Explain it in five or six sentences.
Because autocomplete gets asked on every keystroke, it serves about ten times more traffic than search itself, and each answer must arrive in tens of milliseconds or it loses the race with the next keystroke. A database query for "ca" would fetch millions of matching rows and sort them by popularity just to return ten — and the shortest, most expensive prefixes are exactly the ones everyone types through. No database survives that much work that often. So we flip it: overnight, a batch job computes the ten best completions for every possible prefix and stores them in memory as a giant lookup table — answering a keystroke is one dictionary read. A small streaming job layers on whatever is trending in the last few minutes, so breaking news still shows up. The rule underneath: when the same questions arrive millions of times per second, you stop answering at question time — you precompute the answer sheet.
The PM wins: "suggestions must consider my own recent searches." Add personalization without destroying the CDN hit rate. Sketch the approach and name what you gave up.
Keep the server global and cacheable; personalize at the edge — the client. The device stores the user's own recent searches locally (private data on the user's own hardware — free storage, free privacy win). On each keystroke the client fetches the global top-10 as before (CDN-cacheable, unchanged) and merges: matching local history pinned on top, global results filling the rest. Gave up: personalization is shallow — one device's history, no learned model, no cross-device sync. Kept: every shared cache layer, the entire cost model, and most of the perceived benefit, since users mostly want their own recent searches back. Deep model-driven personalization is a different, far more expensive system — say so with the price tag attached.
Traffic 10×: 50 billion searches/day. Recompute the load and storage, and name the first two components you'd change — and one you deliberately wouldn't touch.
Load: ~600K searches/sec average → ~6M suggest requests/sec, north of 10M peak. Storage barely moves — 10× the searches mostly re-vote for the same queries; unique prefixes might grow 2–3×, say ~10–15 TB → 100–150 shards. First change: edge capacity and short-prefix local copies — the funnel's top absorbs proportionally more, and hit rates improve with traffic density. Second: the stream counter — switch to sampling (1-in-100 is plenty; trending needs signal, not census). Untouched: the core serving design — hash-sharded immutable snapshots in RAM scale linearly by adding shards, which is exactly why we chose it. If your architecture changes shape at 10×, it was the wrong architecture at 1×.
Requirement flip: not web search anymore — autocomplete for an e-commerce catalog of 50 million products, and suggestions must never show out-of-stock items. What changes?
Two things flip. First, scale collapses: 50M products generate a few hundred million prefixes — single-digit gigabytes. The whole table fits in RAM on every node; sharding disappears. Deleting components is the senior move here. Second, freshness hardens from "nice" to "correctness-ish": stock changes minute-to-minute, and a daily build guarantees wrong suggestions. Don't rebuild constantly — decouple: build ranked lists daily (popularity moves slowly), store top-30 per prefix instead of top-10, and filter at read time against a live stock bitmap (50M booleans, a few MB, updated by stock events within seconds), returning the first 10 in-stock. Per-request filtering is affordable now precisely because scale collapsed. Same skeleton, opposite trade-offs — which is why memorizing architectures instead of reasoning from load numbers fails.
Global suggest p99 jumped from 80 ms to 240 ms in an hour. Server-side latency is flat at 4 ms. Your first three checks, in order, with reasoning.
Server-flat-but-user-slow means the regression lives between user and origin. (1) CDN hit rate, per region: the classic cause is an edge cache going cold — a config push changed the cache key (a new query param will do it), a TTL got zeroed, or a POP (point of presence — one physical edge location) went down and rerouted users to distant edges. A regional hit-rate cliff matching the p99 timeline closes the case. (2) A client release ramping that hour: broken debounce (flooding), broken response-cancellation, or a bundle regression shows up exactly as client-measured latency with innocent servers. (3) Response size: new fields or disabled compression spilling into extra round trips moves p99 by exactly this magnitude on mobile RTTs. Meta-lesson: instrument from the client, or you'll page the only team whose component is healthy.
Estimation drill: the batch build must process one day of logs — 5B searches — and emit snapshot files for 50 shards within 4 hours. Roughly how much parallelism, and where's the shuffle?
5B log lines at ~100 bytes ≈ 500 GB/day of raw input — small by batch standards. Two phases: count query frequencies (map over logs, shuffle by query, sum with decay against the historical table), then explode queries into prefixes and shuffle by prefix to pick each top-10. The prefix explosion is the fat middle: 2B kept queries × ~20 prefixes ≈ 40B intermediate records, a few TB shuffled. At ~50 MB/sec of useful shuffle throughput per worker, a few TB in 4 hours needs on the order of 100 workers — a small job, which is the point: the write path is cheap because it's offline. Landing within 5× passes; the graded part is knowing the shuffle-by-prefix is where the bytes are.
It's 8:58am. You're on call for a fintech app with 100 million users. Marketing has scheduled a promo blast — "Flash sale! 40% off gold purchases" — to 40 million phones, going out at 9:00 sharp. At 9:00 the campaign fires. At 9:03, support tickets start arriving. Not about the sale. The tickets say: "I can't log in. My OTP never came."
You open the dashboard. The SMS queue has 3 million messages in it. Somewhere in the middle of that pile, behind hundreds of thousands of promo texts, sit a few thousand two-factor login codes — codes that expire in five minutes, for users who are staring at their screens right now. The users retry. Each retry generates another OTP into the same queue. The pile grows. Your notification system is working exactly as built, and it is locking your own customers out of their money.
When an interviewer says "design a notification system," it sounds like a warm-up — take a message, send it to a phone, done. It is not a warm-up. It's a fan-out problem stapled to a priority problem stapled to a distributed-systems problem where the last hop runs on someone else's servers that throttle you, fail on you, and lie to you. That last part is the interesting part: this is the rare system where the final mile is infrastructure you do not control.
I'd scope it out loud like this. Functional: the system delivers notifications over four channels — mobile push (iOS and Android), SMS, email, and in-app (a bell-icon inbox). Any internal service can be a producer: the auth service sends OTPs, the order service sends "your order shipped," the marketing service sends campaigns. Users control preferences per category and per channel ("order updates: push yes, email no"), including quiet hours. Some notifications are scheduled for later. Out of scope: composing marketing campaigns, audience analytics dashboards, and the details of email content rendering — fair?
Non-functional — and these carry the design:
Here's the trap in estimating this system: counting events. Events are tiny. One marketing campaign is one event. The real unit of work is event × recipients, because every recipient needs their own preference check, their own dedup key, their own device tokens, their own rendered payload.
Transactional traffic: 100M daily users, maybe 5 transactional notifications each per day — 500M sends/day, about 6,000/sec average, call it 20,000/sec at peak. A big number but a smooth one. Now the broadcast: one campaign to 40M users, and product wants it delivered "within 10 minutes" so it lands before the sale starts. 40M ÷ 600 seconds ≈ 70,000 notifications per second — from a single row in a campaigns table. Peak system throughput is therefore ~90–100K/sec through the pipeline, and it arrives as a wall, not a ramp. That number makes the first architectural decision for me: this is fan-out at write, through durable queues, processed by horizontally scaled workers. No synchronous path survives a 40M-recipient write.
Now the number that gates the crux decision. Our SMS providers, combined, drain maybe 1,000 messages/sec (SMS throughput is contractual and carrier-limited — you buy it, you don't scale it). If a campaign ever puts 500K SMS in front of an OTP in a shared queue, that OTP waits 500 seconds. The OTP expires at 300. The user retries, adding more load. So: queue position is latency, drain rates differ per channel by 100x, and therefore priorities must be separate queues with separate workers — not a priority field sorted inside one queue. The math forces the design; we'll commit properly in the deep dive.
Storage: log every send attempt for tracking — ~500 bytes × 1B/day ≈ 500GB/day, ~45TB at 90-day retention. That's an LSM-style store — log-structured, built to swallow writes (Cassandra/Scylla) — partitioned by user_id — straight from the sharding chapter. Device tokens: 100M users × ~1.3 devices = 130M rows. Small. Gate state (dedup keys, rate counters): at 100K sends/sec with ~4 lookups each, that's 400K ops/sec against Redis at burst — fine for a modest cluster, but it tells me gates must be O(1) in-memory checks, never SQL queries.
One internal API for all producers:
POST /v1/notifications
{
"event_id": "order-shipped-88213", // producer-chosen, idempotency root
"priority": "P1",
"category": "order_updates",
"audience": { "user_ids": [...] } | { "segment": "gold-buyers-in" },
"template": "order_shipped_v3",
"data": { "order_id": "...", "eta": "..." },
"channels": ["push", "email"], // omit = per-user preference decides
"ttl_sec": 7200,
"send_at": null // or a future timestamp
}
Core tables: preferences (user_id, category, channel opt-ins, quiet-hours window, timezone), device_tokens (user_id, token, platform, last_seen_at), notification_log (event_id, user_id, channel, status, timestamps — the tracking store), plus an in-app inbox table. Dedup keys and rate counters live in Redis with TTLs — a time-to-live, the expiry after which the store deletes the key for you — not in a database.
Write path: a producer POSTs one event. The ingest API validates it, writes it durably to a raw-events queue, and returns 202 immediately — producers never wait on fan-out. Fan-out workers pick up the event and resolve the audience: for direct sends, the user IDs are in the request; for a segment ("all gold buyers in India"), we read a precomputed member list, not a live database scan — resolving 40M users must be a file read handed out in chunks, not a query. For each recipient, the worker runs the gates pipeline (next section), renders the template into a per-channel payload, and enqueues it into the right queue: one queue per channel per priority class. Channel workers — APNs workers, FCM workers, SMS workers, email workers — consume their queues and talk to the external gateways. In-app is the easy channel: write to the inbox table, push over a websocket if the user is connected.
Read path: three readers. The user reads their in-app inbox (GET /inbox?cursor=…, straight from the inbox table). The producer reads status ("was event X delivered to user Y?") from the tracking store. And the phone-facing channels have a delayed read path of their own: delivery receipts. APNs/FCM tell us about dead tokens; SMS providers send delivery reports (DLRs) minutes later via webhook; email providers report bounces. Those receipts update the tracking store and drive token cleanup and provider health scores. Delivery status is eventually consistent — by design, because the truth arrives late.
This system is a post office, and the crucial detail is that you only run the sorting floor. Letters pour in from thousands of senders. Your job: check the recipient actually wants mail from this sender (preferences), don't deliver at midnight (quiet hours), don't dump forty catalogs on one doorstep in a day (rate limit), and don't deliver the same letter twice when a sender nervously mails a copy (dedup). Then you hand every letter to an outside carrier — the airline that flies it (APNs, Twilio) — and the airline is not yours. It has its own capacity, its own bad days, and it tells you "delivered" with a shrug and a delay. Where the analogy breaks: a real post office has one priority lane. You will need entirely separate trucks.
The hard part of this problem is the last mile you don't control, multiplied by fan-out. Every notification must pass per-user gates (preferences, quiet hours, rate limit, dedup) at 100K/sec, and then be handed to a third-party gateway with at-least-once semantics — because we promised never to lose a P0 — without buzzing the phone twice, even though the gateway gives us no transaction, sometimes no timely acknowledgment, and no way to un-send. Exactly-once delivery to a phone is impossible; the crux is engineering which failure you get — a rare duplicate or a rare loss — and choosing per priority class. I want to spend most of our time here.
Every (event, recipient) pair walks through four gates before it becomes a queued message. Order matters — cheapest and most-likely-to-drop first:
event_id + user_id + channel — the exact pattern from the idempotency chapter, applied per recipient. Check Redis: seen? Drop. This runs first because producer retries are common (the ingest API returned 202 but the producer timed out and re-sent), and it's one O(1) lookup.Then the template render: template + locale + data → a per-channel payload (push JSON, SMS text, email MIME). Render after the gates — rendering 40M payloads for messages that 30% of users opted out of is pure waste.
Now the delivery handshake. A channel worker pulls a message, sends it to APNs, and acks the queue. Two crash windows, and they fail in opposite directions:
Neither "check-then-send" nor "mark-then-send" alone is safe. I'll commit to the lease pattern from the idempotency chapter: before sending, SET key "inflight" NX PX 60000 — claim the key for 60 seconds. If the claim fails and the value is sent, skip (duplicate suppressed). If it's a stale inflight lease, the previous worker died mid-flight — take over and retry. After the gateway accepts, overwrite the key to sent with a 24-hour TTL, then ack the queue.
The honest residue: if the worker sends, then dies before writing sent, and the lease expires, the retry produces a duplicate. That window is real and cannot be closed, because "send to APNs" and "write to Redis" are two systems with no transaction spanning them. So the senior move is to choose the failure direction per priority: for P0, bias toward duplicates — a user who gets two copies of an OTP shrugs; a user who gets zero is locked out. For P2, bias toward loss — a marketing push that quietly disappears costs nothing. Same machinery, different retry aggressiveness, chosen on purpose.
Sending through APNs is dropping a letter through a mail slot in a wall. You can't see through the wall. If the lights go out the moment your hand comes back — did the letter go through? Post it again and maybe two arrive; don't and maybe none did. No amount of cleverness on your side of the wall creates certainty, because the certainty would have to live on both sides at once. All you choose is which mistake you'd rather make — and for a login code, "two letters" beats "no letter" every time.
Back to the OTP stuck behind the campaign. The instinctive fix is a priority field: one queue, sort critical first. In a distributed queue that's somewhere between hard and fake — Kafka partitions are FIFO, and even with a priority-supporting broker, 40M enqueued messages ahead of you means the sort itself is the bottleneck, and one slow bulk consumer still holds shared resources. I'll commit to physical isolation: per channel, separate queues per priority class, separate worker pools, separate provider budgets.
sms.critical drains through dedicated SMS workers with a reserved slice of provider throughput — say 200 of our 1,000 msg/sec is untouchable by anything else.sms.bulk gets the leftovers, rate-capped so it can never consume the reserve. A 40M campaign makes its own queue deep — bulk latency stretches to hours, which the P2 SLA explicitly allows — and the critical lane never notices.The cost, stated plainly: more queues, more worker fleets, more dashboards — roughly 3× the operational surface per channel. Worth it, because the failure it prevents is the worst one this system has: self-inflicted denial of service on login. Head-of-line blocking — one item at the front of a line holding up everything behind it — is a queueing law, not a bug; the only real fix is a second lane.
On June 17, 2021, a large chunk of HBO Max's subscriber base received an email with the subject line "Integration Test Email #1" — an empty template blasted to the real production mailing list. HBO Max's own tweet became famous: "We mistakenly sent out an empty test email to a portion of our HBO Max mailing list this evening… yes, it was the intern." Twitter responded with #DearIntern, thousands of engineers sharing their own worst sends in solidarity. The kind part of the story is the intern being forgiven; the engineering part is why it happened at all: the platform allowed a test event to resolve to a production audience of millions with no guardrail between "send" and forty million inboxes. A notification platform is the last line of defense against every producer's worst day. That's why the design above puts audience-size checks, staged ramp-up (send to 1%, watch, then proceed), and a per-campaign kill switch inside the platform — not in the good intentions of whoever clicks the button.
Everything so far assumed the gateway at the end behaves. It doesn't. Three behaviors to design for:
They throttle you. APNs and FCM are generous but will push back (HTTP 429/503) under bursts; SMS throughput is a hard purchased ceiling; SES enforces sending quotas. Channel workers therefore treat provider capacity as a token bucket — an allowance of sends that refills at a fixed rate — and respect backpressure signals — when the provider says slow down, workers slow down and the queue absorbs the difference. That's what the queue is for. Retries use exponential backoff with jitter (Part 2 reflex — a synchronized retry wave after a provider blip is a self-made DDoS). After N failed attempts, the message goes to a dead-letter queue for inspection, not into an infinite retry loop.
They fail — so expire messages honestly. Every message carries a TTL matched to its meaning. A marketing push has ttl = 4h; if APNs is down for five hours, those messages should die in the queue, unsent (APNs even supports this natively — the apns-expiration header tells Apple to discard undeliverable pushes after a deadline). And here's the counterintuitive one: P0 gets the shortest TTL of all — an OTP older than two minutes must be dropped, because the code has expired, the user has already re-requested, and delivering hour-old login codes in a burst after recovery is confusing at best and an attack surface at worst. Never-lose-a-P0 means "never lose it while it's still meaningful." Retry hard within the window; drop dead-certain at its end.
They lie. A 200 from APNs means "accepted," not "delivered" — the push can still silently die if the device is offline for days. FCM happily accepts messages for tokens whose app was uninstalled weeks ago. SMS providers accept your message and only find out from the carrier minutes later that the number is disconnected — the truth arrives as an async delivery receipt, or never. Consequence one: track sent (we handed it over) and delivered (a receipt confirmed it) as separate statuses in the tracking store, updated asynchronously by the receipts webhook. Consequence two: dashboards and provider-health decisions run on delivered rate, because "sent" is a vanity metric when the provider is lying.
SMS deserves its own reliability answer because any single provider has regional bad days (a carrier route in one country degrades while the provider's API happily returns 200). I'll run two providers minimum, with a router in front: per country-route, track a health score — an exponentially weighted moving average of acceptance rate and DLR (delivery receipt) rate over the last few minutes. Healthy: split traffic by cost, cheapest first. Degraded route: circuit-break it and shift traffic to the backup, automatically, with a manual override for the humans. The cost: two contracts, two webhook formats, message-ID mapping for both, and the failover provider is usually pricier — which is fine, because P0 fails over unconditionally while bulk can simply wait out the blip on the cheap provider.
Push tokens rot. Users uninstall the app, switch phones, or the OS rotates the token — a few percent of your token table dies every month. If you keep sending to dead tokens, two things happen: your delivery-rate metric sinks into noise, and providers notice — a sender with a high invalid-token rate looks like a spammer and can get throttled. So the feedback loop is mandatory: APNs' HTTP/2 API returns 410 Unregistered (with a timestamp) for a dead token; FCM returns UNREGISTERED. The receipts pipeline deletes those tokens immediately. On the front end, the app re-registers its token on every launch and we stamp last_seen_at; tokens silent for 90+ days get demoted out of broadcast audiences. Hygiene, not glamour — and skipping it is a slow-motion outage.
Batching/digests: nobody wants 37 pushes titled "someone liked your photo." For digestible categories, low-priority notifications land in a per-user buffer; flush on a count threshold or a time window ("37 people liked your photo" after 30 minutes), whichever comes first. On-device, set a collapse key (APNs apns-collapse-id, FCM collapse_key) so a newer notification replaces the older one instead of stacking. Digests are P2 by definition and go through every gate.
Scheduled sends (send_at in the API, and quiet-hours deferrals) go to a job scheduler — the design of which is its own chapter (p3c11). One rule matters here: gates run at send time, not schedule time. A user who unsubscribes on Tuesday must not get the blast scheduled on Monday.
The incident pattern to design against is the marketing blast that melts 2FA. It's an entire genre of outage: a bulk producer saturates a shared resource (queue, worker pool, provider quota) and the highest-stakes traffic in the company dies quietly behind it. Priority isolation is the structural fix; the operational fixes on top: campaigns ramp (1% → 10% → 100% with delivery-rate checks between stages), an audience-size threshold above which a send requires a second human approval, and a per-campaign kill switch that stops queued sends — the queue is per-campaign-taggable precisely so a bad blast can be surgically drained.
On January 13, 2018, at 8:07am, phones across Hawaii buzzed with: "BALLISTIC MISSILE THREAT INBOUND TO HAWAII. SEEK IMMEDIATE SHELTER. THIS IS NOT A DRILL." There was no missile. A shift-change drill was running at the Hawaii Emergency Management Agency, and the warning officer on duty missed the "exercise, exercise, exercise" cue on the recorded message, thought the threat was real, and chose the live alert from a drop-down menu where the drill option and the real one sat side by side, separated only by a confirmation click. The FCC's investigation faulted the ambiguous drill and the menu — and found a deeper failure: the correction went out 38 minutes later, because no pre-scripted "false alarm" message existed in the system and officials had to compose and authorize one while the state panicked. The fixes are a notification-platform checklist: two-person confirmation for the highest-stakes sends, test environments that physically cannot reach production audiences, and — the one everyone forgets — a pre-built correction path, because a notification system has no undo. You cannot recall a million buzzes. You can only be fast with the follow-up, and only if you built it before you needed it.
What I'd monitor, per channel × priority: age of the oldest message in each queue (better than depth — a deep bulk queue is normal, an old critical message is an incident), send rate vs. baseline, gateway error rate by code (429s = throttling, 410s spiking = token-table problem, 5xx = provider incident), delivered rate from receipts, and token-invalidation rate. Plus one end-to-end canary: a synthetic user receives a real push through the full pipeline every minute; the alert fires when the canary goes quiet — this catches whole-pipeline failures that every per-component metric misses. Alert thresholds differ by lane: critical-queue oldest-message age > 30s pages someone at 3am; bulk-queue depth alone never does.
Rollout: template and routing changes ship behind flags to 1% of traffic first — a malformed payload can crash a mobile app in the field, and you cannot roll back a million delivered pushes. Provider config changes (new SMS route, new APNs cert/key) get canaried against the synthetic user before real traffic. Certificates and provider keys expire; their expiry dates are monitored like disk space.
What breaks at 10x (1B users, multiple 400M blasts a day): first, audience resolution — "resolve 400M users" must become a precomputed segment file split into thousands of chunks checkpointed by fan-out workers, because any single scanner is hours of work and a single point of restart-from-zero. Second, the gate store — 4M Redis ops/sec needs a properly sharded cluster with per-user hash-tagging so one user's gates stay on one node. Third — and this one has no engineering fix — SMS: provider throughput is bought, not scaled, and at 10x the honest answer is to migrate 2FA off SMS toward TOTP apps (the six-digit codes an authenticator app generates on the phone itself, no network hop) and passkeys, and reserve SMS for the long tail. Sometimes the senior answer to "how does this scale?" is "this channel doesn't — here's the product change that sidesteps it."
Interviewers use this problem to test whether you respect boundaries you don't control, and whether "priority" is a word or an architecture in your answer. Expect these, nearly verbatim:
| Decision | Why I committed to it | The cost I accept |
|---|---|---|
| Fan-out at write, through durable queues | One event → 40M sends; no synchronous path survives that, and queues absorb provider outages for free | Delivery status is async; producers get 202 "accepted," never "delivered" |
| Physical priority isolation (queues + workers + budgets per class) | Queue position is latency; a priority field can't stop head-of-line blocking behind 40M messages | ~3× operational surface per channel: more queues, fleets, dashboards |
| At-least-once + idempotency-key dedup, duplicate-biased for P0 | Exactly-once into a third party is impossible; a lost OTP locks a user out, a doubled one is a shrug | Rare duplicates in the send-then-crash window; Redis is a hot dependency of the send path |
| Per-priority TTLs; expired messages are dropped unsent | An hour-old OTP or a stale flash-sale push is worse than nothing; queues must not become time capsules | During long outages we knowingly discard messages — "never lose" is bounded by "still meaningful" |
| Two SMS providers, health-routed on delivered rate | Single provider = single point of failure we can't fix from outside; accept-rate 200s can't be trusted | Double contracts and integration work; failover route costs more per message |
| Tracking via async receipts; "sent" ≠ "delivered" | The truth about delivery arrives minutes late or never; pretending otherwise poisons every metric | Eventual consistency in status; support sees "sent" for messages that will never arrive |
A junior asks: "Why can't we just guarantee each notification is delivered exactly once? Retry until it works, dedup so it never doubles." Explain in five or six sentences why that guarantee is impossible here, and what we do instead.
Because the last step — handing the message to Apple's or Google's or a carrier's servers — happens outside our system, and there's no transaction that covers both their side and ours. After we send, we might crash before recording that we sent; on restart we can't tell whether the message got through, so we either retry (risking a duplicate) or don't (risking a loss). Recording before sending just flips the failure: crash after recording but before sending, and the message is lost while our books say it went. So instead of chasing an impossible guarantee, we choose which mistake to make, per message type: for login codes we retry aggressively — a duplicate OTP is harmless but a missing one locks someone out — and for marketing we let unclear cases die quietly. Dedup keys shrink the duplicate window to a rare crash-timing accident; they can't close it. Exactly-once isn't a feature you build here — it's a direction you lean.
Add a digest feature: "likes" on a social app should batch into "N people liked your post" instead of N pushes. Design the batching logic — where does the buffer live, what flushes it, and what happens if the user opens the app mid-window?
Per-user, per-post buffer in Redis (key: digest:{user}:{post}, a counter plus first/last actor names for the copy). Flush on whichever comes first: count threshold (say 25) or time window (30 min), driven by a delayed job set on first like. The flushed digest is a normal P2 notification — full gates apply, and it carries a collapse key so a later digest for the same post replaces the earlier one on the device rather than stacking. If the user opens the app mid-window, the in-app read marks the buffer consumed and cancels the pending flush — pushing someone a summary of likes they've already seen on screen is the digest version of a duplicate. Edge to note: the buffer is lossy if Redis dies; acceptable for likes (P2), which is exactly why 2FA never goes near this path.
APNs goes down hard for one hour during your evening peak (20K pushes/sec). Do the queue math: what builds up, what should your system do during the hour, and — the part most candidates miss — what must it NOT do in the first minutes after APNs recovers?
Build-up: 20K/sec × 3600s = 72M queued pushes. During the outage: workers see connection failures/5xx, back off with jitter, and stop hammering; the queue absorbs — that's its job, and 72M × ~1KB ≈ 72GB is nothing for Kafka-class storage. TTLs quietly prune: OTP-class pushes expire in minutes (users already retried via SMS fallback), time-sensitive bulk dies at its 4-hour TTL. On recovery, the trap: 72M messages and a fleet of eager workers are a thundering herd aimed at a service that just got back on its feet — you will re-injure the patient and get throttled into a second outage. Drain through a rate limiter ramping from a fraction of normal throughput upward while watching APNs error rates, critical queues first (they're small — minutes to clear), bulk over the following hours. Bonus signal: mention that during the outage your delivered-rate dashboards go dark, so the on-call needs an explicit "provider outage" status to suppress a wall of misleading alerts.
Requirement flip: this is now a chat app, and push notifications for messages must arrive in order per conversation — a preview of message 5 before message 4 confuses users. What changes in the architecture, and what does ordering cost you?
Ordering requires serializing per conversation: partition the push queue by conversation_id (or user_id) so one partition — consumed by one worker at a time — holds a conversation's messages in sequence, exactly the Kafka partitioning model from the streams chapter. The costs: (1) head-of-line blocking returns in miniature — one undeliverable message stalls that conversation's queue until retry or TTL, so you need a skip-after-N-attempts rule that trades ordering back for progress; (2) hot partitions — a 5,000-member group chat serializes through one partition; (3) worker parallelism is capped by partition count. Cheaper alternative worth saying out loud: don't order the pushes — put a sequence number in the payload and let the client drop stale previews, moving the problem to the edge where it's one integer comparison. Collapse keys (newest replaces oldest) get you most of the UX for none of the server cost. Committing to client-side ordering with a reason is the senior answer; server-side FIFO is the fallback if clients can't be changed.
10x the broadcast: a news app with 500M users wants breaking-news pushes delivered within 5 minutes of the editor hitting send. Which component breaks first, and walk the numbers to your redesign.
Target rate: 500M ÷ 300s ≈ 1.7M sends/sec — 17x our design's burst. First break: audience resolution — a single scanner reading 500M users at even 100K rows/sec needs 83 minutes before the first push leaves. Fix: precompute the all-users audience as a segment file in object storage, pre-split into ~5,000 chunks of 100K users; on send, fan the chunks out to thousands of workers, each checkpointing per chunk so a crashed worker re-does 100K users, not 500M. Second break: the gate store — ~4 lookups × 1.7M/sec ≈ 7M Redis ops/sec, needing a properly sharded cluster; better, for breaking news specifically, collapse per-user gates to a bitmap check (opted-out bitmap in memory on each worker, refreshed every minute) since breaking news ignores rate limits and quiet hours anyway. Third: APNs/FCM connection management — thousands of HTTP/2 connections with tuned concurrent-stream limits, pre-warmed, because opening them during the event is too late. Note what did NOT need redesign: queues, priority lanes, dedup — the skeleton holds; the feeding of it is what changes.
Now add multi-tenancy: your notification system becomes a platform product (like Twilio or OneSignal) serving 2,000 business customers. One tenant queues 80M sends. What's the new failure mode, and what do you add?
The noisy neighbor: priority classes no longer protect you, because the 80M blast and a tiny tenant's P0 traffic can both be legitimately "bulk" or "critical" — isolation must now be per-tenant as well. Add: per-tenant rate limits and quotas enforced at ingest (token bucket per tenant per channel); fair-share scheduling across tenants in the worker layer (weighted round-robin over per-tenant sub-queues, so 80M queued messages from tenant A still leaves tenant B's 50 messages draining immediately); per-tenant provider budgets so one tenant can't exhaust the shared SMS throughput; and per-tenant dashboards/alerts, because their on-call is now your customer support. The deeper shift to name: shared provider reputation. One tenant sending spam gets your shared IPs and sender IDs throttled for everyone — so tenant isolation extends to reputation (dedicated sending identities for big tenants) and an abuse-detection gate at ingest becomes a first-class citizen, not an afterthought.
"Design WhatsApp." The interviewer says it casually, like they're asking you to design a doorbell. And honestly, sending a text from phone A to phone B sounds like a doorbell. Message in, message out. Fifteen minutes, tops. What's the catch?
The catch is the phone. Your user is on a train going through a tunnel. Their connection drops mid-send, comes back for four seconds, drops again, comes back on a different cell tower with a different IP address. Somewhere in that chaos they typed "I'm at the hospital, come now" and hit send. That message is not allowed to vanish. Not allowed to arrive twice. Not allowed to arrive after the message they sent next. A chat app is a promise made over the least reliable network humans use, and the interview is about how you keep that promise.
So say the crux out loud in the first five minutes: "The interesting problem here is guaranteed, ordered delivery over flaky mobile connections — I want to spend most of our time on the acknowledgment protocol and offline delivery." That sentence tells the interviewer you've seen this fire before. Now let's perform the 45 minutes at the senior bar.
WhatsApp is many products wearing one icon. Scope it down, out loud: "I'll build 1:1 chat and small groups — up to 1,024 members, WhatsApp's real cap. In scope: delivery and read receipts, offline delivery, history synced across a user's devices, and online/last-seen presence. Out of scope: calls, stories, payments. Typing indicators I'll cut too — if we have spare time they're a five-minute add-on, and I'll say why they're architecturally free. Fair?"
Functional requirements: send and receive in 1:1 chats and groups; see ticks (sent, delivered, read); receive everything missed while offline; open a new phone and see history; see whether a contact is online.
Non-functional requirements are where the design comes from, so put numbers on them and commit:
Three calculations, each gating a decision — the only reason to do math at a whiteboard.
Message rate. 50 billion over 86,400 seconds ≈ 580K messages per second average. Evenings peak around 3x — call it 1.5–2 million per second, and New Year's Eve blows past even that. Conclusion: no single anything sits on the message path. Every tier shards.
Concurrent connections. The number most candidates miss, and the one that shapes the architecture. Sub-second delivery means no polling — every online user holds a persistent connection. If 40% of a billion users are connected at any moment, that's 400 million open sockets. A well-tuned server holds 1–2 million idle connections (an idle socket costs memory and a heartbeat, not CPU). Take the conservative end: 400M ÷ 1M is 400 gateways to carry the load, ~500 with failure headroom — a tier sized by socket count, not request throughput. A different scaling axis than anything we've designed before — it's why this chapter exists.
Storage. A text message is ~200 bytes with metadata. 50B/day × 200B ≈ 10 TB/day, ~3.6 PB/year before replication. Write-heavy, append-only, never updated, read as "recent messages in one conversation." That shape screams wide-column LSM store (Cassandra/HBase-class — Chapter 1.13), not a relational database. Media is 100x bigger but boring: blob storage behind a CDN (Chapter 1.7), message carries a pointer. Decision made; media never appears again.
Messages ride a WebSocket — a long-lived two-way connection either side can push down at any time. The frames:
client → server: send {client_msg_id, convo_id, payload}
server → client: ack {client_msg_id, seq}
server → client: message {convo_id, seq, sender, payload}
client → server: receipt {convo_id, seq, type: delivered|read}
Plain HTTPS covers the non-realtime edges: GET /messages?convo_id=&before_seq=&limit= for history, POST /media for uploads. Core tables: conversations (members, type), messages (keyed by conversation plus a server-assigned sequence number — much more soon), and inbox — one queue per user per device, holding pointers to messages that device hasn't fetched. Note client_msg_id: the client generates it, and it's about to do a lot of work.
Everything you've built so far had one comforting property: stateless app servers — any request can hit any box. A chat gateway breaks that rule, deliberately. When phone B connects, its socket lives on one specific machine — gateway 42 — and delivering to B means finding that machine. Stateful servers are the exception you accept because the requirement (server pushes to phone, instantly) demands it.
So the machine needs a directory: the connection registry, a sharded Redis keeping user_id, device_id → gateway_id — written on connect, refreshed by heartbeat, with a ~60-second TTL so crashed gateways' entries evaporate on their own. The path of one message:
Walk it aloud: A's send frame arrives at gateway 17, which hands it to the message service. The service durably persists it and assigns a sequence number — then and only then does A get the ack. Next it looks up B in the registry and forwards to gateway 42, which pushes down B's socket. If B isn't in the registry, the message waits in B's inbox and a push notification through FCM/APNs wakes the app (Chapter 3.3's whole subject). Write path and read path both exist; the loop is closed. Time to deep-dive.
In September 2015, WIRED reported a number that made the industry blink: WhatsApp was serving 900 million users with roughly 50 engineers. Facebook, with a similar-sized crowd, had thousands. Half the secret was ruthless scope — no feeds, no ads, just messaging. The engineering half was Erlang on FreeBSD: a language Ericsson built for telephone switches — millions of cheap, isolated, crash-and-restart processes, one per connection, with hot code reloading so you deploy without dropping sockets. WhatsApp engineer Rick Reed had pushed a single tuned server past 2 million concurrent TCP connections in 2012, and the production fleet ran around a million per box. So our "400M sockets ÷ ~1M per box ≈ a few hundred gateways" is roughly the shape WhatsApp actually ran — quoting production history, not guessing.
This is the hard part the question was invented to probe. Delivery over a flaky network isn't made reliable by hoping the network behaves — it's made reliable by a ladder of acknowledgments, each rung a distinct promise from a distinct party, plus retries that are safe to repeat. Every tick in the WhatsApp UI is a rung of a distributed-systems protocol. Candidates who treat ticks as UI polish fail this question; candidates who can say precisely what each tick proves, who sends it, and what happens if it never comes pass at the senior level. Spend your time here.
Rung one: client → server. A's phone sends the message and starts a retry timer. Until the server's ack arrives, the phone keeps the message in local storage and resends — after 2 seconds, 4, 8 — surviving app restarts, tunnels, and tower switches. The single tick ✓ means exactly one thing: the server has this message durably. Which forces the golden rule: persist before you ack. Write to Cassandra (and the recipient's inbox), then ack. Never the reverse. An ack-then-persist server that crashes in the gap has lied: the phone shows ✓, deletes its local copy, and a message the sender was told was safe is gone forever with a checkmark on it.
But retries create a new bug: if the ack gets lost, the phone resends a message the server already has. Duplicate. This is exactly what Chapter 2.2 solved with idempotency keys, and here the key is client_msg_id — a UUID the phone generates at send time. The server keeps a short-lived dedup record (id → seq, TTL a day or so); a resend of a known id gets the same ack back and writes nothing. Retries plus idempotent ids turn flaky networks from a correctness problem into a mere latency problem. That sentence is worth saying verbatim at the whiteboard.
Rung two: server → recipient, confirmed. When B's phone receives the message, it sends a delivered receipt up its own socket; the server records it and pushes ✓✓ to A. The double tick proves not "the server tried" but "B's device has it in local storage." Rung three: B opens the chat, a read receipt flows, A's ticks turn blue. Three rungs, three different promises, from three different places.
The ack ladder is registered post. You hand the parcel over the counter and the clerk stamps your receipt — but only after the parcel is physically in the post office's custody, because the receipt makes them responsible for it. That's the single tick, and why the stamp can't come first. Later a delivery confirmation arrives — the parcel reached the house — the double tick. Whether they've opened it is a third, separate fact — the blue tick. And if you never got a receipt, you assume the post office never got the parcel, and you bring it back: the client retry. One honest gap: real post offices don't handle you bringing the same parcel twice; our dedup ids are the extra machinery for that.
Now the second promise. Suppose we order messages by the phone's clock. A's phone runs 40 seconds fast (phone clocks drift, and users change them). A writes "yes", B writes "wait, no" — and half the group sees them swapped. Worse: a phone coming online sends a message composed an hour ago. Client timestamps are opinions, not facts. Display them; never sort by them.
The fix is the one we already smuggled in: the server assigns each message a per-conversation sequence number at durable-write time. Like the ticket machine at a deli counter — your place in line is decided at the counter, not by when you left home. Order is decided once, centrally; every device sorts by seq forever after, and a client holding 41 and 43 knows it has a gap and asks for 42. Gap-free numbering needs a single writer per conversation: shard conversations across message-service instances by consistent hashing (Chapter 1.9), so one instance owns convo 9 and increments its counter — with a fencing token (Chapter 2.3) so a stale owner can't hand out duplicates during failover.
Notice what we did not build: any ordering across conversations. Per-conversation sequencing means a chatty 1,024-member group funnels through one owner — a deliberate, bounded hot spot, fine because one group's humans can only type so fast. The moment "group" becomes "channel with 5 million subscribers," this breaks — hold that thought for deep dive 3.
B's phone is dead in a drawer. Messages keep arriving. Where do they wait?
In B's inbox: a per-device queue of pointers ("convo 9, seq 42…47"), appended in the same durable write as the message itself. When B reconnects, the gateway drains the inbox in sequence order; B acks the highest seq received; everything at or below is trimmed. Waking a dead phone is the push notification's job — FCM/APNs delivers a contentless "you have messages," the OS launches the app, the app connects and drains. The push is a doorbell, not a mail slot: unreliable, sometimes late, sometimes dropped — and that's fine, because the inbox is the source of truth and the next connect for any reason drains it. Reliability lives in the inbox; the push is just a kick.
The inbox is the pigeonhole rack at a hostel reception. While you're away, mail piles up in your slot in arrival order — reception never runs after you with letters. When you walk in, you empty your pigeonhole in one go, oldest first. The push notification is reception texting "you've got mail": helpful, but nothing breaks if you miss it, because the letters are still in the slot when you show up.
Multi-device is why the inbox is per device. B's phone and laptop each hold their own cursor — the highest seq drained per conversation — so a laptop closed for a week catches up independently while the phone stays live. A brand-new device gets history the other way: paging backwards through the messages table (before_seq) on demand, not replaying 200,000 inbox pointers. Receipts need one rule: "delivered" fires when the first device gets the message, "read" when any device reads it — and read state syncs across devices so the laptop doesn't re-badge chats the phone already read. Undrained inbox rows get a TTL (30 days, WhatsApp's real policy) so an abandoned device can't grow a queue forever.
Commit: Cassandra-class wide-column store. Why: 580K writes/sec is bread and butter for an LSM engine that turns writes into sequential appends; the read pattern ("latest N of one conversation") maps onto one partition read; and messages are immutable, so we give up nothing by giving up transactions. Postgres would need heroic sharding to take this write volume; Cassandra takes it natively. This is Discord's design, near enough: they moved trillions of messages through Cassandra and later ScyllaDB on the same data model.
The trap is the partition key. conversation_id alone means one partition holds a conversation's entire history — and a busy group grows that partition forever. Unbounded partitions are the Chapter 2.6 disease: multi-gigabyte partitions, compaction pain, one hot node. The fix is worth naming precisely: partition by (conversation_id, time_bucket) — say one bucket per month — clustered by seq descending. Every partition is now bounded, "recent messages" reads the newest bucket, and paging history walks backwards bucket by bucket.
Groups reuse the whole machine. A group message is one durable write to the messages table plus fan-out on write to each member's inbox — up to 1,024 pointer appends, batched. And one optimization pays immediately: members' sockets cluster on a few hundred gateways, so route one copy per gateway and let it fan out to its local sockets, not 1,024 copies through the routing tier. But say the boundary out loud: fan-out-on-write scales with recipients, so a broadcast channel with 5 million subscribers means 5 million inbox writes per message. There the model flips to fan-out-on-read — subscribers pull from one channel timeline — the celebrity problem, treated fully in the news feed chapter (3.5).
Presence looks trivial — a green dot — and it's the easiest place to accidentally design a system bigger than the messaging path. The data is simple: the gateway heartbeat we already have (every ~30 seconds) updates last_seen[user] in Redis; online = heartbeat within the last minute. The danger is distribution. A phone on weak signal flaps — connect, drop, connect — several times a minute. Naively broadcasting every transition to ~200 contacts turns one flapping phone into hundreds of updates a minute; at 400M connected users, presence would drown message traffic.
So don't gossip every flicker. Three throttles: subscribe, don't broadcast — push presence only to users with that chat open right now (a handful); everyone else fetches last-seen on demand. Debounce — publish a transition only if it held ~10 seconds, which erases flapping. Degrade first — presence is the designated first thing shed under load (Chapter 2.7 thinking): a stale dot costs nothing, a delayed message costs trust. Privacy rides the read path: "last seen: nobody" just filters the read. And typing indicators, if asked, are presence's little sibling: fire-and-forget frames to open-chat subscribers, no durability, dropped freely under load — which is why cutting them in minute one cost nothing.
Real WhatsApp runs the Signal protocol: messages are encrypted on the sender's device and only recipients' devices hold keys. Architecturally — the honest part — almost nothing we built changes, because every mechanism operates on envelopes, not contents: sequences, acks, inboxes, and receipts never read the payload, and now the payload is ciphertext. What E2E kills is every feature that needed the server to read messages: server-side search, link previews, cloud history for a new device. (Real WhatsApp deletes messages from its servers on delivery, keeps undelivered ones ~30 days, and moves history via encrypted backups — our server-side-history design is the Messenger/Telegram-cloud model, a fork I'd state out loud.) Multi-device gets harder: senders encrypt per recipient device. And metadata — who talks to whom, when — stays visible regardless. In an interview this is the right depth: name what changes, name what doesn't, don't derail into cryptography.
A gateway dies with a million connections. The sockets vanish; a million phones notice within seconds and all try to reconnect. This is the reconnect storm — chat's signature failure mode — and unmanaged, it can hammer the LB, auth, and registry hard enough to knock over healthy gateways: one dead box becomes a cascade. The defenses are boring and mandatory: jittered exponential backoff baked into the client (wait a random 0–5 s, then 2× with jitter — shipped long before the outage, because you can't push code to phones during one); gateways interchangeable at connect time so the LB spreads the refugees; reconnect rate as a first-class metric, your earliest outage alarm. Messages sent to B during the gap? Routing fails, B is treated as offline, inbox plus push catch it. Nothing lost — durability never lived on the gateway.
Deploys use the same muscle deliberately. Never kill a gateway with a million sockets; drain it: stop accepting new connections, then tell existing clients in small random batches to reconnect elsewhere, over ten minutes. A deploy is a controlled, slow-motion version of the failure you already survived — which is exactly why it's safe.
Region failover. Messages and inboxes replicate asynchronously to a second region (Chapter 2.5). Async means an honest RPO: lose a region and the last few seconds of acked writes may not have replicated. I'd commit to async anyway — synchronous cross-region acks put 80–150 ms of speed-of-light tax inside every tick, poisoning the p95 for a once-a-decade event. Then name the safety net: the ack ladder is an end-to-end repair protocol. Any message not yet delivered still sits on the sender's phone with one grey tick; after failover those can be resent, and client_msg_id dedup makes resending safe. The ladder we built for flaky networks quietly covers datacenter loss too.
What I'd watch (Chapter 2.8 discipline): send→delivered p50/p99 for online pairs — the product-truth metric; ack latency (server health); reconnect rate (storm detector); inbox backlog depth and age (is offline delivery draining?); per-gateway connection count; push success rate. Page on delivery p99 and reconnect spikes; dashboard the rest. At 10x: gateways and Cassandra scale by adding boxes, since partitions stay bounded — the real pain is amplification: group fan-out multiplying inbox writes, and the registry under 10x-sized reconnect herds. I'd invest there first.
New Year's Eve is chat's Super Bowl: the whole planet sends "Happy New Year" inside the same rolling midnight hour, time zone by time zone. On December 31, 2017, WhatsApp processed a record 75 billion messages in one day — 13 billion images, 5 billion videos — with about 20 billion from India alone, whose midnight lands in one synchronized burst. The previous record, 63 billion, was set exactly one year earlier: the peak grows double-digit percent every year, on schedule. And on that same New Year's Eve the service went dark for roughly an hour across several countries — WhatsApp never published a cause, so nobody outside the company can pin it on the load, but the timing is a useful piece of theatre. Both halves belong in your interview: chat load is brutally calendar-driven, so capacity-plan for the year's known worst hour rather than the average second, and design the degradation path yourself — because the one night everybody is watching is the worst night to be improvising it.
What the interviewer probes next, nearly verbatim — each aimed at a rung of the ladder or a stateful corner:
| Committed choice | Why | The bill you pay |
|---|---|---|
| Stateful WebSocket gateways | Server-push in <500 ms; polling can't do it at 400M users | Lose stateless scaling; need a registry, draining, and storm defenses |
| Durable write before ack | The ✓ must be a promise, not a guess | Storage write sits inside every send's latency budget |
| Client-generated ids + server dedup | Makes infinite retries safe over flaky links | Dedup table on the hot path; trust clients to generate ids |
| Server-assigned per-convo seq | One order authority; gap detection free; clocks can't lie | Single writer per convo — a deliberate, bounded hot spot |
| Cassandra, key (convo, month-bucket) | Append-heavy load, partition-local reads, bounded partitions | No ad-hoc queries or transactions; bucket-walking history code |
| Fan-out on write to inboxes (≤1,024) | Cheap reads, instant delivery, offline for free | N writes per message; breaks at channel scale — flip to pull there |
| Inbox as truth, push as kick | Reliability where it's controllable; OS push is best-effort | Two paths to build; offline latency depends on OS wake behavior |
| Async cross-region replication | Keeps ~100 ms of speed-of-light out of every ack | Seconds of RPO on region loss — disclosed, partly repaired by the ladder |
Explain to a junior engineer, in five or six sentences, why WhatsApp's server must save a message to disk before sending the single grey tick — and why resending a message can never create duplicates.
The grey tick isn't decoration — it's a promise that the server now owns your message, so your phone is allowed to stop worrying about it. If the server ticked first and saved second, a crash in between would break the promise: your phone shows ✓ and moves on, but the message never reached durable storage and is silently gone. So the order is law: write to storage, then ack. The other half: until the tick arrives, your phone re-sends the message again and again, and re-sends could create duplicates — except every message carries an id your phone generated when you pressed send. The server remembers ids it has already stored, so a re-send of a known id just gets the same ack again and writes nothing. Saving-before-acking means nothing accepted is ever lost; ids-with-dedup mean retrying is always safe — together they turn a terrible network into a mere annoyance.
Add typing indicators to the design in three sentences. What guarantees do they need, and what infrastructure do they reuse?
When you type, your phone sends a fire-and-forget frame through your gateway, routed only to users who currently have this chat open — the presence subscription list we already maintain. No durability, no inbox, no receipts, no retries: if it's lost, the next keystroke sends another, and under load these frames are dropped first. Typing indicators are presence with a shorter fuse — they reuse the sockets, routing, and subscriptions the real system already built, which is why cutting them from scope cost nothing.
It's December 31st. Modeling from WhatsApp's real numbers (75B messages that day in 2017, growing yearly), your peak hour will run 4–5x normal peak, and you can't buy 5x hardware. What degrades, in what order, and what must never degrade?
Never degrades: the durable write and the ack — the promise is the product. Shed in order: (1) presence and typing entirely; (2) batch receipts — coalesce delivered/read updates, seconds late instead of instant; (3) defer media processing — accept uploads but delay thumbnailing, letting texts jump the queue; (4) let inbox drains and group fan-out lag by seconds, since inboxes absorb backlog by design. Every step converts "instant" cosmetics into "seconds late" while accepted-means-durable stays absolute. And pre-warm — NYE is the one peak you can circle on a calendar, so capacity-test against last year's record times 1.3, because it grows every year.
Requirement flip: enterprise demands server-side search over message history ("find every message mentioning invoice #4412"). What does this cost, and how would you build the search path?
First cost, said honestly: it kills end-to-end encryption — the server can't index what it can't read, so this is now the Slack/Messenger trust model, a product decision to surface, not sneak past. Build-wise: don't query Cassandra — wide-column stores can't text-search. Emit every stored message onto a stream (CDC into Kafka, Chapter 2.9), consume into an inverted index (Elasticsearch-class, Chapter 3.19) partitioned per user or workspace so searches stay partition-local. Search is eventually consistent — findable seconds after sending, stated as an explicit SLA. The messages table stays the source of truth; the index is a rebuildable derived view, so index corruption is an inconvenience, not data loss.
Estimate the gateway fleet: 300M concurrent connections, boxes with 64 GB RAM, roughly 20 KB of memory per idle connection. Don't forget failure headroom.
Per box: budget ~40 GB for connection state (rest for OS, TLS, spikes) → 40 GB / 20 KB ≈ 2M theoretical; run at half for burst safety → 1M per box. 300M / 1M = 300 boxes minimum. Headroom: a dead box's 1M connections must land on survivors during a reconnect storm, and an AZ loss moves tens of boxes' worth at once — provision ~30% spare: ~400 gateways. Sanity-check: WhatsApp ran 1–2M connections per tuned FreeBSD/Erlang box, so the per-box number is historically grounded. Notice CPU never entered the math — idle sockets cost memory and heartbeats, which is why this tier sizes on RAM and file descriptors.
Product ships "Communities": broadcast channels where one admin posts to 5M subscribers. Your group machinery does 1,024-member fan-out-on-write. What breaks, and what's the redesign?
Fan-out-on-write breaks arithmetically: one post = 5M inbox writes; a few popular channels posting daily and inbox writes dwarf all real messaging. The sequencer is fine (one admin writes), but delivery flips to fan-out-on-read: the channel gets a single time-bucketed timeline; subscribers pull new posts on app-open, nudged by one batched push. Online subscribers still get a cheap nudge via the per-gateway multicast trick — one frame per gateway, not per user. Receipts change meaning: no per-user ticks (5M ticks per post is its own storm) — approximate view counts instead (Chapter 2.10 sketches). This is the push→pull boundary in miniature; the full treatment is the news feed chapter (3.5).
"Design the Facebook news feed." At Meta, this is the house specialty. There's a real chance your interviewer spent years working on the actual feed, which means two things: they will not be impressed by buzzwords, and they know exactly which room is on fire. One more thing before the clock starts — Meta runs this question in two flavors (Chapter 0.2). In the System Design round, the backend machine is the star. In the Product Architecture round — the flavor most product engineers get — your API contract and client behavior are graded too: what's in the payload, what the cursor promises, what the app does while a request is in flight. This walkthrough plays it as a Product Architecture candidate performing at E5, because that's the harder grading.
Here's the trap hiding inside the question. You can design a feed that works beautifully for 499 million of your 500 million daily users — and it will melt for the last few thousand. Those few thousand are the celebrities, and the biggest of them — Cristiano Ronaldo, on Meta's other feed, Instagram — has over 600 million followers. A single tap of his thumb generates more write traffic than the rest of the platform combined for the next several minutes.
The whole interview is secretly about that one moment. Everything else — caching, ranking, counters — is important supporting cast. Let's run the 45 minutes.
I scope out loud and get a nod before drawing anything. Functional: users follow other users (asymmetric — I follow you, you don't follow me back); users create posts with text and media; users see a home feed of posts from accounts they follow, ranked by relevance, not just time; users like and comment, and everyone sees the counts. Out of scope, and I say so: ads, Stories, the notification system (that's Chapter 3.3), and comment threads beyond a count.
Non-functional, with numbers I'm committing to:
Notice the CAP posture (Chapter 1.10) fell out of the requirements for free: the feed chooses availability and takes eventual consistency on freshness and counts. A stale feed is Tuesday. A down feed is a headline.
I do only the math that changes decisions (Chapter 1.4). First number: the read/write ratio. 500M daily users open the app a few times a day and scroll a few pages each time — call it 10 feed-page fetches per user per day. That's 5 billion fetches a day, about 60K per second average, so I plan for 200K reads/sec at peak. Writes: maybe one in five daily users posts, so 100M posts a day — about 1.2K/sec, call it 5K/sec at peak. That's roughly 50 reads for every write. Conclusion, said out loud: I will happily do extra work at write time to make reads a cheap cache hit. That is the argument for precomputing feeds — pushing each new post into followers' feeds when it's created, so a feed read is just "get my list."
Second number: the fanout multiplication. Fanout is the copying step — one post fanning out into many follower feeds — so pushing a post means one write per follower. An average account has ~200 followers, so 1.2K posts/sec becomes ~240K tiny cache writes/sec, maybe a million at peak. A sharded Redis fleet shrugs at that; each write is a few microseconds of list-push. So far, push looks great.
Then I do the celebrity row of the table, because follower counts are a power law (Chapter 2.6) — almost everyone has a few hundred followers, a handful of accounts have tens of millions, and the average is a lie told by the tail: a 100M-follower account posts once, and that's 100 million writes for one tap. At a generous 1M writes/sec of fanout capacity, one post occupies the entire fleet for 100 seconds. A Ronaldo-sized Instagram account at 600M followers: ten minutes of everything we own, and every ordinary user's post queued behind it. Also: 100M copies of a post ID at ~20 bytes each is 2 GB of RAM duplicated per celebrity post — most of it for followers who won't even open the app today. This is the number that forces a hybrid design, and I say so now: "push for normal accounts, pull for celebrities — I'll deep-dive the threshold later."
One more sizing check, because it decides what the feed cache stores: 500M active users × 300 feed entries × 16 bytes per entry (post ID plus a timestamp) ≈ 2.4 TB of RAM — about 40 Redis shards. Perfectly buildable, but only because I'm storing post IDs, not post bodies, and only for active users. Store full posts and it's hundreds of terabytes. The estimate just made two design decisions for me.
This is the part the Product Architecture flavor of the round grades hardest (Chapter 0.2), so I'm precise but fast:
POST /v1/media → returns an upload URL for blob storage and a media_id. The client uploads bytes directly to blob storage, not through my app servers (Chapter 1.7).POST /v1/posts with {text, media_ids} → 201 with the created post.GET /v1/feed?cursor=…&limit=25 → {items: […], next_cursor}. The cursor is opaque — clients must not parse it, so I can change its encoding without breaking apps in the field.PUT /v1/posts/{id}/like and DELETE …/like — PUT because liking is naturally idempotent (Chapter 2.2): double-tap, retry, same result.POST /v1/follows with {followee_id}.Each feed item comes back fully hydrated: post body, author name and avatar, like/comment counts, CDN URLs for media. One request paints the screen. And I name the client behaviors Meta grades in the Product Architecture flavor: render the cached last page instantly on app open while fetching fresh behind it; show a "new posts" pill instead of yanking the scroll position; bump the like count optimistically on tap and reconcile later; prefetch the next page when the user nears the bottom.
Data model: a users table; a follows table that I need in both directions — "who do I follow" (for building feeds) and "who follows X" (for fanout) — so it's stored twice, sharded by each key (Chapter 1.9); a posts table sharded by author_id with time-sortable IDs (Chapter 1.14), which means "recent posts by author X" is one cheap indexed read. Hold that thought — that per-author recent-posts index is called the author's outbox, and it's about to matter enormously.
Write path first, because the write path is where the interesting decision lives.
Walking it: the post service validates, writes the post to the sharded posts DB — durability first, this write is the promise behind the 201 — then drops a small "post created" event on a durable queue (Chapter 1.11) and returns. The user sees their own post immediately (more on that trick in the ops section). Fanout workers consume events, ask the social graph service for the follower list, and pipeline the post ID into each follower's Redis feed list — push to the front, trim the list to ~300 entries. Workers checkpoint queue offsets, so a crashed worker replays; pushes are idempotent-enough because a duplicate ID gets dropped at read by the seen-filter. Two cheap optimizations I name unprompted: skip followers who haven't opened the app in 30 days (a huge fraction of any follower list — they'll rebuild by pull if they return), and batch pushes per Redis shard so one post's fanout is a few hundred pipelined calls, not 200 individual round trips.
Now the read path — the one that runs 200K times a second.
Gather candidate post IDs: the user's precomputed list (a Redis hit), plus the recent outboxes of any mega-accounts they follow, plus their own recent posts. Filter out already-seen IDs. Score the ~500 survivors (deep dive 2). Take the top 25 and hydrate them — post bodies and author profiles from a read-through post cache (Chapter 1.6; post bodies are immutable, which makes them a caching dream), like/comment counts from the counter service, media as CDN URLs. Those three hydration fetches run in parallel, and each one is allowed to fail without killing the page: if the counter service is slow, ship the feed with slightly stale counts rather than hold 25 posts hostage for a number nobody will fact-check.
Hydration is a thali counter. The plate (your feed page) gets assembled from separate stations — rice station, dal station, pickle station — all spooning in parallel, not one queue through all three. And if the pickle station is backed up, the server hands you the thali without pickle instead of letting your whole meal go cold. That's parallel fetch with graceful degradation. Where the analogy breaks: a thali missing dal is a complaint; a feed missing a like-count for three seconds is something no user has ever noticed.
What if the Redis list isn't there — new user, evicted user, dead shard? Fall back to pull: read the follow list, fetch each followee's outbox, merge by time, rank, serve, and repopulate the cache in the background. Slower — dozens of outbox reads instead of one list read — but correct, because the posts DB is always the source of truth. The same pull machinery handles a new follow (backfill the followee's recent posts, or let the next read's merge pick them up) and a dormant user returning after months. One mechanism, four problems solved — and I say so out loud.
Push or pull is not a style preference — it's forced by the shape of the follow graph, and the follow graph is a power law. Pull-for-everyone dies on reads: 200K feed requests/sec × ~400 followed accounts each = 80 million outbox reads per second, most of them recomputing the same merge they computed five minutes ago. Push-for-everyone dies on writes: one celebrity post = 100M+ writes, minutes of total fleet capacity for one tap. Neither pure strategy survives, so the design must be hybrid — and the senior signal is where you put the threshold and why. This is the room that's on fire. If you take one thing from this chapter into the interview, take this paragraph.
Push is newspaper home delivery: the work happens at dawn, one copy dropped at every subscriber's door, and reading is instant. Pull is the newsstand: nothing is delivered anywhere; you walk over and ask for today's papers, and they assemble your bundle while you wait. Home delivery is wonderful — until one "paper" has 600 million subscribers and the delivery van fleet spends all day on a single edition. So you do what cities actually do: home delivery for the neighborhood papers, and for the mega-publication everyone reads, you just pick it up at the stand — it's one stand, it's always stocked, and the trip is cheap because you only visit a handful of stands. The analogy breaks in one place: real newsstands have queues, but a celebrity's outbox is a tiny, heavily cached read that scales to any number of readers.
My commit: push for accounts under 1 million followers, pull-and-merge for accounts at or above it. Why 1M? Three curves cross near there. First, fanout time: 1M pushes at ~50K pipelined writes/sec per worker group finishes in ~20 seconds — the top edge of my "followers see it in seconds" budget. Beyond that, delivery time grows linearly into minutes. Second, population: accounts over 1M followers number in the low thousands globally — a rounding error of authors, which is exactly why exempting them is cheap. Third, read cost: a typical user follows maybe 5–20 such accounts, so the read-time merge adds a handful of small, hot, easily cached outbox reads — bounded and predictable. Anywhere from 100K to 5M is defensible; what's graded is that I know both cost curves and picked a point on purpose. The threshold is checked at post time against the author's current follower count, so an account that crosses 1M mid-week just switches paths on its next post — old pushed posts still sit in follower lists, new ones arrive via merge, and nobody notices.
The canonical public account of this exact design is Raffi Krikorian's 2012 "Timelines at Scale" talk, given when he ran Twitter's platform engineering. The numbers he shared are the whole argument in miniature: about 150M active users, roughly 300K timeline reads per second against only ~4K new tweets per second — reads dominating writes by nearly two orders of magnitude. So Twitter precomputed: every home timeline lived in a Redis cluster as a list capped at ~800 entries, replicated three ways, and a fanout service pushed each new tweet into follower timelines in pipelined batches of about 4,000 destinations at a time. Then the power law showed up. Lady Gaga had ~31 million followers, and a single tweet from her took minutes to fan out — long enough that followers sometimes saw replies to her tweet before the tweet itself, because the reply (from a small account) fanned out in seconds while the original was still crawling through 31 million list-pushes. Twitter's fix is the one in this chapter: stop fanning out the mega-accounts, keep their tweets in a separate path, and merge them into timelines at read time. One architecture, forced by one distribution.
That reply-before-the-original race is worth ten seconds in the interview: fanout means different followers receive the same post at different times, so ordering across users is never guaranteed — only ordering within one user's assembled page, which the read path fixes when it merges and ranks.
Meta probes ranking on this question almost every time, and the senior (E5) move is to be honest about the boundary: I don't design the model — I design the system the model lives in, which means candidate generation, a latency budget, feature freshness, and a fallback. That answer scores better than bluffing about neural architectures, because it's the answer of someone who'd actually operate this.
Candidate generation is the merge we already built — pushed list plus celebrity outboxes, a few hundred IDs. Cheap filters run first (seen, muted, blocked authors), because filtering before scoring means paying the expensive step on fewer items. Scoring is a separate service with a hard budget: 50 ms of my 200 ms page budget to score ~500 candidates in one batched call. The model consumes features — how often you interact with this author, predicted probability you'll like or comment, the post's age, whether it has video — served from a feature store kept fresh by the stream pipeline (Chapter 2.9). And here's the probe Meta loves: feature freshness is a silent failure mode. If the pipeline lags a day, the scorer doesn't error — it confidently ranks with stale affinity data, and feed quality quietly rots while every dashboard stays green. So I monitor feature age with the same seriousness as latency, and I say that unprompted.
The fallback: if the scorer misses its deadline or its error rate spikes, the feed service sorts candidates by time and ships that (the degradation menu from Chapter 2.7). Users get a slightly dumber feed, not a spinner. I track chrono_served_rate as a first-class metric: near zero normally, and its climb is my early warning — often before the scorer team's own alerts fire. After scoring, diversity rules re-rank: no five-in-a-row from one author, mix media types, and a freshness floor so the model's love of proven old bangers never fully buries today's posts.
Ranking a feed is as much a product rollout problem as a systems one, and Instagram's 2016 switch is the documented proof. Until then the feed was strictly reverse-chronological. In March 2016 Instagram announced it would reorder by relevance, publishing the motivating number: people were missing, on average, 70 percent of the posts in their feed. The backlash was immediate and loud: users organized around keeping the chronological feed, and celebrities and businesses panic-posted "turn on notifications for our account" for weeks, afraid the algorithm would bury them. Instagram rolled it out anyway — gradually, cohort by cohort, through 2016, measuring as it went. Two years later it reported that ranked feeds had users seeing 90 percent of posts from friends, up from around 50. And in 2022, after years of user and even regulatory pressure, Instagram added chronological options back ("Following" and "Favorites") alongside the ranked default. Every beat of that story maps to a system decision: gradual rollout behind cohorts, guardrail metrics, and keeping the chronological path alive — which happens to be exactly the fallback path your scorer outage needs anyway.
Like counts are their own scaling problem, because likes dwarf posts: 100M posts a day might collect a billion likes — 12K/sec average, with vicious spikes when something goes viral, all aimed at a single post's counter — a textbook hot row (Chapter 2.6). So counts live in a dedicated counter service: increments land on a queue or in sharded in-memory counters, get aggregated, and flush to durable storage in batches. Reads come from a cache that's a few seconds stale. The contract is approximate now, exact eventually — and the product design covers for it beautifully: a viral post displays "1.2M likes," a rounding that hides more drift than the counter service will ever produce. Meanwhile the one number a user will fact-check — whether their own like registered — is handled by the client bumping the count optimistically the instant they tap. That pairing — eventual consistency in the backend, read-your-own-writes faked at the edge — is a Product Architecture answer, and interviewers notice it.
Pagination has a trap: the ranked order changes between requests (new posts arrive, scores shift), so "give me items 26–50" is meaningless — page two would overlap page one. My fix: the first feed request materializes a ranked session snapshot — the scored ID list, stored server-side with a short TTL — and the cursor encodes (session, position). Paging walks the frozen snapshot: stable, duplicate-free, cheap. A pull-to-refresh starts a fresh session, and the seen store — a small per-user set of recently served post IDs, TTL of a few days — keeps the new session from re-serving what the old one already showed. A few hundred IDs per user; at this scale it's Redis pocket change, and it's the honest answer to "no duplicates across sessions."
Cold starts, quickly, since our pull path already does the heavy lifting: a brand-new user with zero follows gets a fallback candidate source (onboarding interests, popular content) — the system design just needs that pluggable source to exist. A new follow backfills the followee's recent posts into your list, or lazily lets the next read's merge find them. A returning dormant user's evicted list rebuilds by pull — one slower first page, then back on the fast path.
Now the section that separates "good design" from "design I'd carry a pager for" — raised before the interviewer asks.
chrono_served_rate pins to 1.0, engagement guardrails sag, and model rollback is a config flag — because every model change ships behind a flag to 1% of users first. Before even that, it runs in shadow mode: it scores real traffic and its output is thrown away, so you learn its latency and its ranking without a single user seeing it.The dashboard, in priority order: feed p99 (the SLO, Chapter 2.8), feed-list cache hit rate (target ~99%; a slide means the pull path is quietly becoming the main path), fanout lag p99, chrono_served_rate, scorer p99, feature age, counter queue depth. At 10x: fanout scales linearly (partitions and workers — embarrassingly parallel), the feed cache goes from 2.4 TB to 24 TB of RAM (more shards behind consistent hashing — though I'd revisit entries-per-list before buying that RAM), and the real money pit is the scorer fleet, where the lever is candidate count: score 200 instead of 500 and the bill more than halves, for a quality loss the guardrails can measure. I might also drop the celebrity threshold, trading more read-time merging for less write amplification. Knowing which knob to turn first is the 10x answer.
How Meta actually probes this design once it's on the board — every one of these has a committed answer earlier in this chapter:
| Choice | Why | The cost I accepted |
|---|---|---|
| Hybrid fanout, threshold 1M followers | Pure push melts on celebrities; pure pull melts on 80M outbox reads/sec | Two code paths, and a read-time merge that must stay fast forever |
| Feed cache stores post IDs only | 2.4 TB of RAM instead of hundreds; bodies cached once, not per-follower | Every read pays a hydration step — mitigated by the post cache |
| Ranked feed with chronological fallback | Relevance drives the product; availability beats cleverness | Occasional "dumb" feeds, and a scorer dependency to budget and monitor |
| Eventually consistent counters | A billion likes/day can't be synchronous row updates | Counts lag seconds and drift until reconciled; hidden by display rounding + optimistic client |
| Session-snapshot cursors + seen store | Ranked order shifts between requests; offsets would duplicate | Server-side session state with a TTL, and a per-user seen set to store |
| Skip fanout to 30-day-inactive followers | Cuts fanout writes massively for followers who won't look | Returning users pay one slow pull-rebuilt page |
| Media via CDN, uploaded direct to blob storage | Feed servers never touch image bytes (Chapter 1.7) | An upload handshake in the API, and cache-invalidation lives with the CDN |
Explain to a junior engineer, in five or six sentences, why the news feed uses push for most accounts but pull for celebrities — and why neither alone works.
Feeds are read about fifty times more often than posts are written, so we do the work at write time: when you post, we push your post's ID into each follower's cached feed list, making a feed read a single cheap lookup. That breaks for celebrities — one post from an account with 100 million followers would mean 100 million writes, tying up the whole fanout system for minutes over one tap. Doing it the other way for everyone — building each feed at read time from every followed account's recent posts — breaks too, because 200K feed reads per second times hundreds of followed accounts is tens of millions of repeated lookups every second. So we split by follower count: under a million followers, push; over it, we skip fanout and just stitch those few accounts' recent posts into your feed when you open the app. It works because almost nobody is a celebrity, and each of us follows only a handful of them — pushing is cheap for the many, pulling is cheap for the few.
Now add Stories: 24-hour-expiry posts with a per-viewer "watched" state, shown in a tray above the feed. Which parts of this chapter's machinery do you reuse, and what's genuinely new?
Reuse almost everything: media through the same blob/CDN pipeline, the outbox pattern (a story is a post with a TTL), and the celebrity split still applies. The tray is easier than the feed — it lists authors with active stories, not items, so pull works fine: fetch your follow list's active-story flags, heavily cached. Genuinely new: expiry (TTL on storage and caches, plus a sweep for the durable copy) and watched-state — like the seen store but per-story-per-viewer and product-visible (the ring dims), so it must be read-your-writes for the viewer: a small write on every view, at feed-read scale. That write volume is the sneaky cost to call out.
World Cup final. An account with 500M followers posts at the final whistle, at the same moment overall posting spikes 10x. Walk both paths and name what melts first, second, and third.
The celebrity post never touches fanout — it lands in the author's outbox and is done; readers pull it at read time, and that outbox becomes the hottest key on the planet, absorbed by its cache and replicas. What melts first is the read side: everyone opens the app at once, so feed QPS spikes well past 200K — the scorer fleet hits its budget ceiling first, and chrono_served_rate climbing is the designed response, not a failure. Second: the counter service, as tens of millions of likes hit one post ID — sharded counters and display rounding are carrying the day. Third: fanout lag for ordinary posts, because the 10x posting spike multiplies through average follower counts; the queue absorbs it and posts arrive late, shedding dormant-follower pushes to catch up. Nothing is lost anywhere — the design bends in exactly the three places it was built to bend.
Requirement flip: a regulator (or a product pivot) demands a strictly chronological, complete feed — every post from every followed account, in order, no ranking. What gets simpler? What gets harder?
Simpler: the scorer, feature store, and diversity layer disappear — with them the 50 ms budget, freshness monitoring, and the fallback (chrono is the feed now). Pagination eases too: time order pages naturally by timestamp cursor, no session snapshot. Harder, and this surprises people: completeness becomes testable. The ranked feed quietly forgave a lost fanout write — the scorer just showed something else, and nobody could prove a post was missing. Now every fanout write matters, trimming lists at 300 entries breaks the promise for prolific-follow users (you must page into outboxes past the cache), and there's no ranking layer to cap candidates. Ranking, it turns out, was also an error-hiding blanket.
Estimation drill: 1B daily actives, feed lists of 800 entries, 24 bytes per entry. How much feed-cache RAM, roughly how many shards, and what else grows that people forget?
1B × 800 × 24 bytes ≈ 19.2 TB — call it 20 TB. At ~100 GB usable per node, ~200 shards before replicas, ~400 with one replica each: a serious fleet with real consistent-hashing weight (Chapter 1.9). The forgotten growers: the seen store (1B users × hundreds of IDs — the same order of magnitude as the feed lists), session snapshots, and fanout throughput, which scales with posts × average followers — several million list-writes per second at peak. Worth saying before buying 200 nodes: challenge the 800. If 300 entries served the product fine, that single number is 60% of the fleet.
Add "delete post." The post may already sit in up to 100M feed caches, session snapshots, and client caches. Design deletion, and say what you explicitly won't do.
What I won't do: chase the copies. Fanning out a delete to 100M lists has the exact celebrity-write problem we built the hybrid to avoid, and it still can't reach client caches. Instead: tombstone the post in the posts DB (the single source of truth) — a tombstone is just a marker row saying "this ID is dead," kept because an absent row and a deleted row must look different to every reader — purge the post cache entry, and let hydration enforce it — every read path already turns IDs into bodies through the post store, so a tombstoned ID hydrates to nothing and is silently dropped from the page. Stale IDs then age out of feed lists via trim and TTL on their own. Client caches get the same treatment on next refresh; a truly urgent takedown can also purge CDN media. The principle worth saying out loud: with many cheap copies of a reference, enforce deletion at the one choke point every read passes through, not at every copy.
Friday, 11pm. A stadium concert lets out and eight thousand people reach for their phones at once. Two riders — Asha and Ben — tap "Request" within the same 200 milliseconds. The nearest driver, Dev, idles two streets away. Two separate matching servers, each handling one request, query the map and find the same answer: Dev. Both send him the trip. Dev's phone shows one; the other rider watches a car icon drive toward them, then vanish. "Driver cancelled." One-star review, refund ticket, a rider who reinstalls a competitor's app.
Now the interview version. The interviewer says: "Design Uber — rider requests a ride, we find a nearby driver, they take the trip." Most candidates hear a geography problem and spend forty minutes on maps. The geography matters — you do need to find nearby drivers in milliseconds out of a million moving dots — but it's the warm-up act. The question was invented to probe the moment above: two riders competing for one driver, and a system that must never give him to both. That's a concurrency problem wearing a geography costume.
Let's walk it the way an E5 candidate should: numbers first, because one number in this problem is so violent it dictates half the architecture before we've drawn a single box.
"I'll scope to the core marketplace loop: rider requests, system finds nearby drivers, ranks them, matches exactly one, driver accepts, we track the trip to completion — plus a lightweight ETA and a touch of surge, since both fall out of the same data. Out of scope: payments internals, turn-by-turn routing, pool rides, food delivery — fair?"
Functional:
Non-functional — with numbers I'll defend:
That last bullet is the thesis of the whole design, and I'll keep returning to it: this system has two kinds of data with opposite personalities.
Driver location pings are the firehose. 1M online drivers, one ping every 4 seconds:
1,000,000 drivers ÷ 4s = 250,000 location writes per second, sustained.
Give each ping a durable database write and you're provisioning a serious sharded cluster to store values that are worthless in 4 seconds, when the next ping replaces them. Meanwhile the data that actually matters is tiny: latest position per driver is maybe 100 bytes (id, lat, lon, cell, status, heading, timestamp). One million drivers × 100 bytes = 100 MB. The entire live map of every Uber driver on Earth fits in the RAM of a laptop.
Now the other side. 20M rides/day ≈ 230 requests per second on average — a few thousand at peak, clustered by city rush hours. That's small. A single Postgres could handle the trip writes for the whole planet.
So the estimation hands us the architecture split on a plate:
Matching's challenge is not throughput at 230/sec. It's correctness under concurrency. That's where the deep-dive time goes.
POST /rides — pickup, destination; an idempotency key header, because a rider on flaky stadium wifi will retry and must not create two requests (Part 2's idempotency chapter, applied).GET /rides/{id} — state + driver position for live tracking.POST /drivers/location — batched pings over a persistent connection.POST /offers/{id}/accept and /decline — driver's response, carrying the offer token (this token becomes load-bearing later).Durable data (Postgres): trips (id, rider_id, driver_id, state, pickup, destination, timestamps, surge multiplier) plus riders and drivers. Ephemeral data (sharded memory): per-driver state {driver_id → lat, lon, cell, status, updated_at} and per-cell sets {cell → set of available driver_ids}. Location history goes to a stream — we'll get there.
Write path (locations): the driver app sends a ping every 4 seconds over a persistent connection. The gateway batches; the location service computes the driver's map cell and updates two in-memory structures on a sharded cluster: the driver's own state record, and the set of driver ids in that cell. Every ping is also appended to Kafka for the analytics world — the hot path never waits for that.
Read/match path (rides): the request creates a durable trip row first (state REQUESTED) — if everything downstream dies, we still know this rider is owed a car. Then matching: query the geo index for candidates near the pickup, rank by road ETA, reserve the top candidate atomically, push the offer, and on accept write the durable transition. Decline or timeout releases the reservation and cascades to candidate #2.
Now the two deep dives that decide the interview: how the map query works, and how the reservation can't be won twice.
The naive approach: store lat and lon columns, index them, query a bounding box. Here's why the database B-tree betrays you. A B-tree is a one-dimensional sorted list. Index on (lat, lon) and it sorts all drivers by latitude first. Your query — "lat between 19.05 and 19.09 AND lon between 72.82 and 72.86" — uses the index for latitude only, then scans every driver on Earth inside that latitude band: a 4-kilometre-tall ribbon wrapped around the entire planet, through Mumbai, Jeddah, and Mexico City. Two dimensions folded into one lose their neighborliness. And the write side is worse: 250K position updates per second churning a durable B-tree. Wrong tool twice over.
The standard fix is to chop the world into named cells and index by cell name. The question that separates candidates is which cells.
Geohash interleaves the bits of latitude and longitude into one string (like te7ud); longer strings mean smaller rectangles, and cells sharing a prefix are near each other. Simple, and it works. But it has three warts. The reverse isn't guaranteed: two points a metre apart can straddle a boundary and share almost no prefix, so nearby queries must always fetch neighbor cells — and computing a rectangle's 8 neighbors is fiddly. The rectangles aren't uniform: their real-world width shrinks away from the equator, so a cell in Oslo covers less area than one in Singapore — annoying when cell statistics feed pricing. And square neighbors are unequal: 4 edge-neighbors close, 4 corner-neighbors ~40% farther, biasing any "look one ring out" logic.
Uber built and open-sourced H3, a hierarchical hexagonal grid. Hexagons fix the square grid's warts: every hexagon has 6 neighbors, all sharing a full edge, all with centers the same distance away; cells at a given resolution are near-uniform in area anywhere on the planet. H3 has 16 resolutions; at resolution 8 a cell averages ~0.74 km² (edge ~460 m), at 9 ~0.1 km² (edge ~175 m). I'll commit to resolution 9 for the driver index — a cell holds tens of drivers in a dense city, not thousands — and a coarser resolution for surge statistics. Honest costs: hexagons don't nest perfectly into parents (cross-resolution containment is approximate), and the grid needs 12 pentagons to wrap a sphere — parked over oceans, ignorable here.
Cell indexing is how you already think about "near me": postal codes. To find pharmacies near your flat you don't sort every pharmacy on Earth by latitude — you check your PIN code, plus the adjacent PIN codes in case you live on the boundary. H3 is a postal system where the government drew every district as an identical honeycomb cell, so "adjacent districts" is always exactly six, in every city, at every scale. Where the analogy breaks: real postal codes follow roads and rivers; H3 cells ignore them — which is exactly why raw cell distance can lie, and why we'll re-rank by road ETA in a moment.
Uber open-sourced H3 in 2018, and the engineering blog is refreshingly blunt about the motivation: it wasn't matching, it was marketplace math. Uber computes supply, demand, and surge pricing per cell across whole cities, and those per-cell numbers constantly get smoothed with their neighbors' numbers. On a square grid, smoothing is biased — diagonal neighbors are farther than edge neighbors, so the same algorithm behaves differently along different compass directions. On triangles it's worse (three neighbor distances). Hexagons are the only regular tiling where every neighbor is equidistant, so a gradient of surge pricing spreads evenly in all directions. Uber accepted a real cost for that property — hexagons can't perfectly subdivide into child hexagons, so H3's hierarchy is approximate, unlike quadtree-style squares that nest exactly — and judged the trade worth it. Choosing a worse hierarchy to get better neighbors is a lovely example of knowing which property your workload actually leans on. H3 now runs far beyond Uber, in mapping and analytics stacks industry-wide.
The structure is almost embarrassingly simple: a hash map cell_id → set(driver_ids) plus driver_id → state, sharded across a memory cluster (Redis or a custom store) by consistent hashing on the key — Part 1's sharding chapter, verbatim. A ping updates in place: same cell, just overwrite lat/lon/timestamp in the driver's record — no set churn. Crossed a boundary? Remove from the old cell's set, add to the new one, atomically (a tiny Lua script on the shard). At res 9 a city driver crosses a cell about once a minute, so perhaps one ping in ten touches a set; the common case is free.
The rider-side query: compute the pickup's cell, then fetch driver sets for that cell plus its k-ring of neighbors (k=1 is 7 cells; k=2 is 19). The ring is not optional: a rider near a cell edge may literally see a driver 80 metres away — across the line — while the nearest driver in their own cell is 400 metres the other way. Query only your own cell and you match the wrong car. If the ring comes back thin (3am, suburb), widen k until you have enough candidates or hit a radius cap.
Straight-line ("crow-flies") distance lies in cities. A driver 200 metres away across a river or a divided highway can be 12 driving minutes away; a driver 1.5 km along your same road is 3. So the ~20 candidates from the k-ring go to an ETA service that answers "driving minutes to this pickup" from the road network and live traffic. I treat it as a black box with a contract: a batch of 20 pairs, answered in under ~100 ms, mildly stale traffic acceptable. It's genuinely a separate system — map data, routing graphs, ML on traffic — and saying "separate problem, here's the interface and my latency budget" is the right scoping move, not a dodge. The pattern: geo index narrows a million to twenty cheaply; ETA spends real compute on only those twenty. Filter cheap, rank expensive, in that order.
Ranking produced "Dev is the best driver for Asha." Meanwhile another matcher concluded "Dev is the best driver for Ben." Both conclusions are correct — that's the trap. Read-then-act is a race: both matchers read Dev as available, then both act on that stale truth. No amount of clever querying fixes this, because the query and the dispatch are separate steps with a gap between them, and at concert-surge concurrency, something always lands in the gap. The fix is the same shape as every race fix in this book: make check-and-claim a single atomic operation at one authoritative place — a per-driver compare-and-set, available → reserved, that exactly one matcher can win. Whoever loses doesn't retry Dev; they take their ranked list and move to candidate #2. The interview is won or lost on whether you walk this race concretely, not on whether you say the word "lock."
Mechanically, Dev's state record lives on exactly one shard of the memory cluster — the single authority for Dev. The reservation is one atomic operation there: "if status is available, set reserved, owner = this match attempt, expiring in 20 seconds" — in Redis terms SET driver:dev reserved NX PX 20000, or a small Lua script (atomic because the shard runs scripts one at a time). Two matchers race; the shard serializes them; one gets true, one gets false. The loser moves down its candidate list — no waiting, no retry storm on the same driver.
Three details make this production-grade rather than whiteboard-grade:
An old hotel front desk has one wooden key rack — one hook per room. Two clerks can both believe room 14 is free (they both glanced at the rack a second ago), but only one hand physically grabs the key; the other clerk's hand closes on an empty hook, shrugs, and reaches for room 15's key. Nobody compares notes, nobody argues — the rack is the single authority and grabbing is atomic. The 20-second TTL is the house rule that if a clerk takes a key and wanders off without checking anyone in, the manager puts the key back on the hook. Where the analogy breaks: our "rack" is sharded across machines — but each key still lives on exactly one hook on exactly one shard, which is the property that makes the grab atomic.
Why not do reservations as row locks in the trips database? Because the reservation spans human decision time — 15 seconds of a driver thinking — and pinning row locks across that, thousands of times a minute in a hot city, turns a database into a queue. The database records outcomes; the memory tier arbitrates contention. One nuance said aloud: Redis-style reservations are best-effort under failover — a dying shard can drop one. That's acceptable, because the blast radius is a duplicate offer (a driver sees two requests, accepts one; the fencing check still prevents a double trip). The durable transition is the true commit point. We're not building a bank on Redis; we're building a turnstile in front of a bank.
Everything so far was allowed to be fast and forgetful. That budget ends the moment Dev taps accept, because from here on, the system's memory is the business: who is driving whom, what will be charged, what the insurance covers. The trip is a durable state machine in Postgres:
Transitions are idempotent and validated ("accept is legal only from OFFERED with a matching token") so a retried accept or double-tapped cancel can't corrupt the story. If every candidate declines, the durable REQUESTED row lets us retry honestly or apologize honestly, instead of losing the rider's request into the void.
Location history rides the same split. The hot path keeps only "latest position"; every ping was also tee'd to Kafka, and consumers batch it into cheap storage. That history is where ETA models train, where "my driver took a weird route" support cases get answered, where safety features replay a trip. Streams glue the fast world to the thorough world (Part 2): the matcher never waits for analytics, and analytics never touches the matcher.
Surge falls out of data we already have. Per coarse hex cell, take the ratio of open requests to available drivers over a sliding window; it drives a price multiplier. Smooth it twice before it touches a price: over time (so one taxi-less minute doesn't spike fares) and over space (blend each cell with its six equidistant neighbors — the exact operation hexagons were chosen for), so surge is a gradient, not a cliff where crossing a street halves your fare. The multiplier is stamped onto the trip row at quote time — durable, because it's money.
Matching server crashes mid-offer. The TTL releases Dev in 20 seconds; the trip row still says OFFERED with no accept, so a timeout sweep re-runs matching. Nothing human-visible except a slightly slower match.
A geo-index shard dies. That slice of drivers vanishes from the map — and reappears within one ping interval, because every driver re-announces themselves every 4 seconds. This is the elegant property of the design, worth saying explicitly in the interview: the geo index is a cache of the world, and the world keeps talking. No backups, no careful replication — the ping stream rebuilds it. Region failover works the same way: point driver connections at the surviving region and the map self-heals in seconds. Derived state that regenerates from its source is state you don't babysit.
In 2015, Uber's chief systems architect Matt Ranney gave a QCon talk on rebuilding the dispatch system ("DISCO") that is still one of the best public looks inside this problem. The geo index was sharded across dispatch nodes using Ringpop — an open-source library Uber built that combines SWIM gossip (nodes constantly whisper membership and health to each other) with consistent hashing, so any node could receive any request and forward it to the node owning that key; supply was indexed by geo cell (Google's S2 cells then — H3 came later). But the most striking idea in the talk was the failover story: rather than synchronously replicating active-trip state across datacenters, Uber leaned on a distributed store it already operated — the driver phones. The app kept an encrypted copy of its own trip state, refreshed with each server exchange; if a datacenter fell over, drivers reconnected to another one and their phones re-uploaded the trips the new datacenter had never seen. The fleet itself was the backup. It's the same principle as our self-healing geo index taken one level further: when the edge of your system holds the truth and keeps repeating it, the middle can afford to forget.
What I'd monitor (the on-call dashboard, unprompted):
Rollout: matching changes ship in shadow mode first — the new ranker runs alongside the old, choices logged, not dispatched — then city by city behind a flag. Cities are natural blast-radius containers; no matching change should ever launch to the whole planet at once.
At 10x: 2.5M pings/sec. The index still fits in ~1 GB; the problem is throughput, so: more shards, harder batching at the gateway, and — the biggest lever — adaptive ping rates: a driver parked at the airport pings every 30 seconds, a downtown driver every 2. Matching QPS stays embarrassingly small, and per-city isolation contains the hot hours. The design bends without snapping.
This problem is a favorite because the geography lures candidates away from the concurrency. The interviewer lets you build the geo index, then aims every follow-up at the gap between "found a driver" and "dispatched a driver":
A passing senior answer walks the race concretely, names the atomic primitive and where it lives, knows why the TTL and token exist, points at the exact boundary where ephemeral becomes durable — and says the self-healing insight out loud instead of proposing to back up a cache.
| Decision | Why | Cost accepted |
|---|---|---|
| Latest locations in sharded memory, no durable write per ping | 250K writes/sec of 4-second-lifetime data; whole dataset is ~100 MB | Shard loss blanks part of the map for a few seconds until pings repopulate |
| H3 hexagons, res 9, over geohash | Uniform-area cells, six equidistant neighbors, clean k-ring queries and surge smoothing | Approximate parent/child nesting; a less familiar dependency than plain geohash strings |
| Per-driver CAS reservation with TTL + fencing token | Makes check-and-claim atomic at one authority; crash-safe by expiry; stale accepts rejected | Best-effort under shard failover — rare duplicate offers possible (never duplicate trips) |
| Sequential offer cascade (one driver at a time) | No driver ever races another for a trip he was shown; simple correctness story | Worst-case match latency stacks decline windows; mitigated by ranking well and short offer windows |
| ETA service as a black box with a 100 ms batch budget | Road-network ranking without dragging routing into this design | Matching quality is hostage to another team's latency and accuracy SLOs |
| Durable Postgres state machine from REQUESTED onward | Money, safety, and support all replay this record; transitions idempotent | An extra durable write before matching starts, on a path where 20 ms doesn't matter |
A junior asks: "Why can't we just query the drivers table for the nearest driver and assign him? Why all this Redis-and-reservation machinery?" Explain in five or six sentences.
Two separate reasons, and they're both about scale of a different kind. First, drivers report their positions 250,000 times a second, and each position is stale in four seconds — writing that firehose into a durable table buys you a huge database bill for data that's worthless almost immediately, when the entire live map fits in 100 MB of memory. Second, and more dangerous: "query then assign" is two steps, and in the gap between them another server can query the same driver and assign him too — now one driver has two riders. So we make claiming a driver a single atomic operation on the one shard that owns his record: check he's available and mark him reserved in one indivisible step, with a timeout so a crashed server can't leave him stuck. Whoever loses that race just moves to the next-best driver. The database still stores every trip durably — it's only the split-second contention that gets settled in memory.
Now add UberPool: riders can share a car, and the system should batch compatible riders heading the same direction. What breaks in your matching design, and what's the minimal structural change?
Sequential greedy matching breaks: one rider's best assignment now depends on other riders' requests, so matching becomes an optimization over a batch. The structural change: hold requests in a short per-area window (a few seconds), then solve jointly — who shares which car, in what pickup order, under per-rider detour limits. The reservation machinery survives unchanged (the solver's output is still "claim these drivers atomically"), but latency is now traded for match quality. DoorDash's dispatch team has written about exactly this: slightly delaying dispatch to batch, and timing assignment against food prep, beat instant greedy assignment.
10x the pings: your fleet grows to 10M concurrent drivers (2.5M pings/sec) but your memory-cluster budget only doubles. Rank your three biggest levers and justify the order.
(1) Adaptive ping rates — the biggest lever, because most drivers at any moment are parked or cruising where 4-second freshness is waste; 30s parked / 8-10s highway cuts ingest 3-5x with near-zero matching impact. (2) Gateway batching and delta encoding — a batch of 50 pings costs little more than one, multiplying shard capacity without touching semantics. (3) Only then, more shards — it works (cells hash cleanly) but it's the expensive lever. Not on the list: durable pings or disk — the workload's nature didn't change, only its volume.
Requirement flip: product wants broadcast offers — show the request to the 5 nearest drivers at once, first to accept wins. Your CEO loves it ("faster matches!"). What does this do to your reservation design, and what do you warn about?
The reservation relocates from before the offer to at the accept: all 5 accepts race, and the atomic CAS now arbitrates accepts — first valid token wins the trip, the other four get "already taken." Same primitive, new position, and it genuinely can cut latency. Warn about the human cost: four drivers did the work of noticing and tapping for nothing — repeated losses train drivers to ignore offers or insta-accept without judgment, degrading the marketplace. The trip row also becomes a contention point, fine at 5, not at 50. A product trade-off wearing an engineering costume; the senior answer quantifies both sides before implementing.
What breaks if you index drivers at H3 resolution 6 (edge ~3.2 km) instead of 9? And at resolution 12 (edge ~9 m)? Name the failure mode of each extreme.
Res 6: huge cells, so a downtown cell holds thousands of drivers — the "narrow cheaply" step stops narrowing, every query drags a giant set into ETA ranking, and busy cells become hot keys (one stadium cell = one shard eating a city's traffic). Res 12: cells smaller than a car, so nearly every ping crosses a boundary (constant set churn; update-in-place dies) and a pickup radius needs k-rings of hundreds of cells. Cell size trades query fan-out against candidate-set size and churn; res 8-9 sits where a cell holds tens of drivers and ~7-19 cells cover a sensible radius.
Estimation drill: your surge system recomputes supply/demand per res-8 hex every 10 seconds. A big city covers ~1,000 active hexes; you operate in 500 cities. Is this computation a scaling problem? Do the math, then say what actually deserves the engineering attention.
500 cities × ~1,000 hexes = 500K hexes, recomputed every 10s = 50K cell-updates/sec, each a trivial ratio plus 6-neighbor smoothing — comfortably one modest service. Compute isn't the problem. The attention belongs on the number's behavior: smoothing windows (too twitchy → price oscillation as drivers chase surge and destroy it; too smooth → surge lags the spike it exists to fix) and boundary effects. The senior instinct: when the arithmetic says "small," redirect effort to correctness and dynamics, not throughput.
"Design Dropbox." The interviewer draws a laptop, a phone, and a desktop on the whiteboard, connects them with a cloud, and adds one sentence that changes everything: "The user edits the same file on two of these while offline. Neither edit may be lost."
Most candidates hear "Dropbox" and think storage problem — big files, lots of disks, a CDN somewhere. That framing fails the interview. Storing bytes is the easy half; blob storage is a solved, buyable product. The question was invented to probe the other half: sync correctness. A student writes two thesis paragraphs on a plane with no Wi-Fi. Earlier that day she fixed a citation on the library desktop. Both machines come online tonight. If your design picks a winner silently, one of those edits evaporates and she finds out the week the thesis is due.
So I'll say the crux out loud in the first two minutes: the hard part of Dropbox is not storing files, it's making independent copies of a file converge without losing work — and doing it cheaply enough that a one-line edit to a 1 GB file doesn't re-upload a gigabyte. Chunking, dedup, and conflict handling. That's where I'll spend the deep dives.
Functional — I'll scope to five things and cut the rest out loud:
Out of scope: real-time co-editing (that's Google Docs, a different machine — p2c04), search, previews, and mobile camera backup (exercise at the end).
Non-functional — with numbers I'll commit to:
Number one: file sizes follow a power law. Most files are tiny — documents, code, configs, a few hundred KB at most. Most bytes live in a few huge files — videos, design files, disk images. This one fact splits the system in two. The small files mean metadata operations (what changed? which version? who can see it?) dominate request volume, so metadata must live in its own fast, strongly consistent service. The huge files mean bandwidth is dominated by big blobs, so the byte-moving path must never re-send what the other side already has. Two workloads, two planes. That's the architecture, derived before drawing a single box.
Number two: the dedup ratio — the number that pays for all the chunking complexity. 500M users × 8 GB ≈ 4 exabytes of logical data. Users store the same things: the same installers, the same shared team folders on ten laptops, twelve near-identical versions of the same presentation. If splitting files into chunks and storing each unique chunk once cuts stored bytes by even a third, that's over an exabyte we never buy disks for. At a rough all-in cost of $100 per usable TB per year (disks, servers, power, replication overhead), an exabyte avoided is on the order of $100M per year. Chunking adds real complexity — hash indexes, refcounts, a garbage collector with a race condition we'll meet later. This number is why the complexity is worth it. If dedup saved 2%, I'd store whole files and go home early.
And a worked example for delta sync, because it gates a design choice. Take a 1 GB slide deck, cut into 4 MB blocks — 256 blocks. You edit three slides; the changed bytes fall in two blocks. Delta sync uploads 8 MB instead of 1 GB. On a 2 Mbps home uplink, 1 GB takes about 67 minutes; 8 MB takes about 32 seconds. That's the difference between "sync is invisible" and "sync is a progress bar you watch." The bandwidth math, not elegance, is why every file becomes a list of chunks.
Four calls carry the whole system. Every file lives in a namespace — one per user, plus one per shared folder.
POST /ns/{ns}/check body: list of block hashes -> which are missing
PUT (presigned URL) upload one 4 MB block to the blob store
POST /ns/{ns}/commit body: path, base_version, ordered block hashes
-> new version, or 409 CONFLICT
GET /ns/{ns}/changes?cursor=C long-poll -> changes since C, new cursor
Data model, five tables: namespaces; files (namespace, path, current version, pointer to manifest); manifests — immutable, an ordered list of (block hash, size) pairs, one per file version; blocks (hash → storage location, size, refcount); journal (namespace, sequence number, change record) — an append-only log of every change in a namespace, which is what cursors read. Sharing is just membership: mount a shared namespace into your view, ACL checked at the namespace boundary.
A file here is IKEA furniture. The manifest is the instruction sheet: "this wardrobe = part 4711, part 8003, part 4711 again…" The blob store is the warehouse, which keeps exactly one bin per part number no matter how many wardrobes use that part. Editing a file doesn't rebuild the wardrobe — it prints a new instruction sheet that reuses most of the old part numbers and adds a few new ones. And "which version is current" is just which instruction sheet is clipped to the front of the folder. Swapping that clip is one tiny atomic action, no matter how big the wardrobe is.
The spine of this design is the split the estimation demanded: a metadata plane and a data plane.
The metadata plane is a sharded SQL database plus a thin API. It holds namespaces, file rows, manifests, ACLs, and the journal. Everything in it is small and must be strongly consistent — "which version is current" is exactly the kind of question where stale answers destroy data. The data plane is a blob store holding 4 MB blocks keyed by their SHA-256 hash. Blocks are immutable and content-addressed: the name of a block is the hash of its content, so a block can never be "updated," only referenced or not. Immutable content-addressed data can be cached anywhere, replicated lazily, and stored on the cheapest thing that spins — it needs none of the coordination the metadata needs. Clients upload and download blocks directly against the blob store using short-lived presigned URLs (p1c07), so the metadata plane never touches bulk bytes.
The write path, end to end. The client watches the filesystem, sees thesis.docx change, and: (1) cuts the file into 4 MB blocks and hashes each one locally; (2) calls check with the hash list — the server answers "I already have 254 of these 256, send me 2"; (3) uploads the 2 missing blocks to the blob store, in parallel, resumable per block; (4) calls commit with the manifest — the ordered hash list — and the version it based its edit on. The server validates every referenced block exists, then atomically advances the file row from base version to new version and appends to the journal.
Step 4 is the linearization point, and I want to name that explicitly. A linearization point is the single instant where an operation goes from "hasn't happened" to "happened" for every observer. Blocks can be half-uploaded, retried, duplicated, abandoned — none of it matters, none of it is visible. The universe only changes when the manifest commit lands in the metadata DB. Before it: old version, everywhere. After it: new version, everywhere. This one property gives us atomic updates of arbitrarily large files using a boring single-row transaction, and it makes every failure before the commit harmless by construction.
The read/sync path. Every device keeps one cursor per namespace — its position in that namespace's journal. Devices hold a long-poll connection to the notifier; when a commit lands, the notifier pokes every device subscribed to that namespace, and each device pulls changes?cursor=C, gets the new manifests, diffs them against local state, downloads only the blocks it doesn't already have, assembles the new file version in a temp file, and renames it into place — the local filesystem's own atomic swap, mirroring the server's. Cursors make sync stateless-server and resumable: a laptop closed for three weeks just replays the journal from its old cursor. Same protocol whether you were gone three seconds or three weeks.
The metadata/data split isn't academic — it's what let Dropbox pull off one of the boldest infrastructure moves of the decade. From its founding, Dropbox kept file blocks in Amazon S3 and metadata on its own servers. Around 2013 they decided the economics no longer worked at their scale and spent two and a half years building Magic Pocket, a custom exabyte-scale blob store, then quietly moved more than 90% of over 500 petabytes of user data onto their own hardware, finishing in early 2016. The migration ran a custom data mover for months while the sync product kept working — possible only because blocks are immutable and content-addressed, so the data plane could be swapped underneath the sync logic like changing warehouses without reprinting a single instruction sheet. Their 2018 IPO filing put the payoff on paper: infrastructure costs fell $39.5M in 2016 and another $35.1M in 2017 — about $75M over two years — and gross margin climbed from 33% in 2015 to 54% in 2016 to 67% in 2017, with the filing naming the migration as a main cause. When your interviewer asks why you're splitting metadata from data, this is the answer with a dollar sign on it.
First decision: how do you cut files into blocks? The obvious way is a ruler: every 4 MB, cut. Fixed-size chunking is fast, simple, and has one famous weakness: insertions shift everything. Add one sentence to page 1 of a 500-page document and every byte after it moves; every 4 MB window now contains different bytes; every hash changes; you re-upload the whole file to change one sentence.
The fix is content-defined chunking: let the file's own content decide where the cuts go. Slide a small window (say 48 bytes) along the file computing a rolling hash — a hash that's cheap to update as the window slides one byte, rather than recomputed from scratch. Whenever the hash's last 22 bits are all zero — which happens on average every ~4 MB, at positions determined purely by local content — declare a boundary. Now insert a sentence on page 1: the bytes shifted, but the same content patterns still produce boundaries at the same content, so a few chunks near the edit change and every later chunk keeps its old boundaries and old hashes. Insertions stay local.
Fixed chunking is tearing a novel into 10-page bundles by page number; content-defined chunking is tearing it at chapter breaks. Add a paragraph to chapter 1 and the page-number bundles all shift — bundle 7 now starts mid-sentence somewhere new, so every bundle looks "changed." The chapter-break bundles don't care: chapter 1 got longer, chapters 2 through 40 are identical bundles you already have. The analogy's honest edge: rolling-hash boundaries are statistical, not guaranteed — you enforce minimum and maximum chunk sizes so a weird file can't produce a million tiny chunks or one giant one.
My commitment: fixed 4 MB blocks, not content-defined chunking — and here's the reasoning. Most real edits are in-place overwrites or appends (office docs rewritten by apps, logs, exports), where fixed blocks already localize the change; insert-heavy byte-shifting edits are the minority. Fixed blocks buy me trivially predictable offsets (block i = bytes at 4i MB — resumable ranged downloads for free), simpler server logic, and less client CPU. This is also what Dropbox actually shipped: 4 MB blocks, SHA-256 per block. The senior part of the commitment is naming the metric that would flip it: I'll dashboard the fleet-wide dedup ratio and bytes-uploaded-per-edit, and if telemetry shows a large population of insert-heavy files paying full re-uploads, content-defined chunking is the targeted upgrade — it changes only the client cutter and nothing in the server contract, because the server never cared where the cuts came from, only what the hashes are.
Second decision, and it's a security decision wearing a storage costume: dedup across users? The chunk hash is the chunk's identity, so dedup falls out naturally — the check call skips any block the store already has. But have it skip blocks that other users uploaded and you've quietly turned a hash into a bearer token: anyone who learns a file's block hashes can "sync" the file into their account without ever possessing it, and the check response doubles as an oracle for "does anyone on this service store this file?"
This isn't hypothetical; it happened to Dropbox. In April 2011, developer Wladimir van der Laan released Dropship, an open-source tool that exploited exactly this design. Dropbox at the time deduplicated across all users: if any user had ever uploaded a block, the client protocol let you claim it by hash alone. Dropship serialized a file's block hashes into a small JSON file; anyone could feed that JSON to the tool and the full file would materialize in their own Dropbox — a movie "transferred" by passing around a kilobyte of hashes, with Dropbox's storage doing all the work. It effectively turned the dedup system into a file-sharing network. Dropbox scrambled to shut it down — takedown notices went to mirrors (Dropbox later said the DMCA notice was fired by an automated banned-file system), and they changed the backend so hash-only claims no longer worked. The lesson is precise: a content hash proves you know a file's name, not that you ever had its bytes. The moment "server already has it" becomes "you may have it," dedup is an access-control bug.
My commitment: scope client-visible dedup to the account, and dedup across accounts only inside the storage layer. Concretely: check answers "already have it" only if the block is already referenced by your account or a namespace you're a member of; otherwise you upload it. Server-side, the blob store still keys blocks by hash globally, so if a thousand accounts upload the same installer it's stored once — the full storage saving survives. What I give up is the bandwidth saving on the first upload of a popular file per account, and I'll pay it gladly: the storage dollars were the big prize from the estimation, and no hash ever acts as a read capability across a trust boundary. (Proof-of-possession challenges — "hash bytes 1 M–2 M with this salt" — can recover some cross-user bandwidth savings later; that's an optimization with a threat model attached, not a day-one feature.)
Everything so far is plumbing a strong mid-level candidate can assemble. The reason this question exists is here: two devices edit the same file while disconnected, and the system must converge without human-invisible data loss. There is no clever algorithm that merges two arbitrary binary edits correctly — a .docx is a zip file; "merging" two zips byte-wise produces garbage. So the design goal is not to be smart. It is to detect concurrency reliably, refuse to guess, and preserve both versions where the human can see them. The mechanism is a base-version check at the linearization point plus the conflicted-copy pattern. Get this right and the rest of the chapter is engineering; get it wrong and you built a data-loss machine with great bandwidth efficiency.
Walk the scenario precisely. thesis.docx is at version 7 everywhere. Laptop A (on a plane) edits offline. Desktop B (at a library desk) edits offline. Both hold the same fact: "my edit is based on v7." A lands first: commits with base_version 7, server sees current = 7, accepts, current becomes 8. B commits with base_version 7 — but current is now 8. The base-version check is a compare-and-swap: "advance 7→my-version only if still 7." It fails. That failure is the concurrency detector: B's edit provably did not see A's edit, so these are concurrent, not sequential.
Now the committed policy: never last-write-wins, never silent merge — fork. B's client receives the 409, downloads v8, and commits its own local content as a new file: thesis (B's conflicted copy 2026-08-14).docx. Both versions now exist side by side in the folder on every device; the human — the only agent that understands thesis semantics — merges. This is Dropbox's actual documented behavior, right down to the filename format with the device name and date baked in so you can tell at a glance which machine and when.
Why commit so hard against LWW? Because LWW answers "who wins?" with a clock, and clocks don't know anything about homework. Whichever laptop syncs second — not edits second, syncs second — silently annihilates the other's work, and the victim gets no signal until they go looking for a paragraph that no longer exists. A conflicted copy is ugly; users mock the filename. But ugly-and-loud beats clean-and-lossy in any system whose one promise is "never lose your files." The conflict is a fact about reality — two humans really did diverge — and the system's job is to report facts, not hide them.
One honest simplification to flag to the interviewer: I'm using a plain base-version integer, not a version vector. Version vectors — one counter per device, so you can order any two versions or prove them concurrent — are necessary when there's no single authority, e.g. peer-to-peer sync or multi-master replication. Here the metadata shard is the single serializer for its namespace, so "does current still equal base?" answers the only question that matters, with one integer. Simpler machine, same guarantee — as long as the metadata DB stays single-writer per namespace, which I've made a load-bearing property and will name again under failure modes.
Edge policies, stated fast because the interviewer will probe them: edit-vs-delete conflict → the edit wins and the file resurrects (deleting content someone is actively editing is the destructive guess; trash makes real deletes recoverable anyway). Concurrent renames → one wins by CAS, the other becomes a second name conflict-copied the same way. Two devices creating the same path independently → first commit wins, second forks. One policy generates all of these: when in doubt, keep both; never destroy content on a guess.
How hard is this state machine really? Dropbox's original sync engine ("Sync Engine Classic," ~2007) accumulated a decade of edge cases — shared folders, selective sync, case-insensitive filesystems, offline forks — until its own engineers described it as nearly impossible to reason about or extend. In 2020 they shipped Nucleus, a ground-up rewrite in Rust, and the published design centers on exactly the concept this deep dive is built on: the client keeps three trees — the Remote Tree (server's latest state), the Local Tree (what's on disk), and the Synced Tree (the last state both sides agreed on — the recorded base). Diff Remote against Synced: that's what the server changed. Diff Local against Synced: that's what you changed. Both differ for the same node: that's a conflict, detected structurally rather than guessed. They validated the engine with randomized simulation testing — millions of generated offline/online interleavings replayed against the state machine, hunting for any sequence that loses a file. When a company rewrites the heart of a product serving hundreds of millions of users mostly to make this logic provable, believe them about where the crux is.
Version history and trash cost almost nothing to build in this design, and that's worth saying in the interview: manifests are immutable, so every version of every file is already a frozen snapshot — a list of hashes. "Restore Tuesday's version" = commit Tuesday's manifest as the new current. Trash = keep the file row, flagged, for 30 days. No copies made, ever, because blocks are shared between versions by reference.
The bill arrives as growth: old versions pin blocks forever unless something deletes them. So: each block carries a refcount — how many live manifests reference it. Prune a version past retention → decrement its blocks. Refcount zero → the block is garbage. Delete it, reclaim an exabyte-scale storage bill's worth of dead bytes over time. Simple.
Except refcount-zero-then-delete has a race that causes silent data loss, and spotting it unprompted is a strong senior signal. Recall the upload flow: check says "block h already exists, don't send it." The client obediently doesn't. Between that answer and the client's commit — seconds, or hours if the laptop lids shut — the last other reference to h gets pruned, refcount hits zero, GC deletes h. The commit then lands, validates… or worse, doesn't validate, and now a live manifest points at bytes that no longer exist. The file corrupts on next download, discovered weeks later.
The fix is a lease plus a grace period — fencing logic (p2c03) wearing a GC hat. A check answer of "exists" writes a short-lived lease pinning that hash (say 24 hours, matching the upload session TTL); commit renews or consumes it. And GC never deletes at the moment refcount hits zero: it marks the block a candidate, waits a grace window strictly longer than any lease can live, re-verifies the count and lease table, then deletes. Two-phase, deliberately slow. The deep principle to say out loud: in storage systems, deleting late is an efficiency bug; deleting early is data loss. Every tie breaks toward late. The same grace-window sweep also mops up orphan blocks from crashed uploads — one mechanism, both messes.
What breaks, and why it's mostly fine by construction:
check; only still-missing blocks travel. Orphans → GC.changes — latency grows from seconds to minutes, nothing is lost, cursors don't care. Graceful degradation was designed in, not bolted on.The scaling bottleneck is the metadata DB, and I'd say so before being asked. Blob storage scales flat — immutable content-addressed blocks shard perfectly by hash. The metadata plane is where the pressure lands: 100M daily devices × checks, commits, cursor pulls. Shard by namespace_id: every row a commit touches lives on one shard, so the CAS transaction never crosses shards, and a shared folder — its own namespace — lands wholly on one shard too. The hot-partition risk (p2c06) is the 10,000-employee company's shared namespace: every commit fans out to 10,000 cursors. Mitigations in order: cache journal tails aggressively (fan-out is read-heavy and identical for all listeners), batch notifications (coalesce N changes in a second into one poke), and only then consider splitting monster namespaces.
At 10x (a billion devices): the notifier's held connections and the journal fan-out melt first — that tier goes cell-based (p2c05). Metadata shards split further; namespace-sharding already gave us the clean cut lines. The blob store mostly just buys more disks. And dedup ratio becomes a board-level number: at 40 EB logical, each percentage point of dedup is real money.
What I'd watch on the dashboard — the four numbers that describe this system's health: sync lag (commit → other-device-applied, p95 — the product promise in one metric); conflict rate (a step change means a client bug is mis-detecting bases, not that users suddenly got collaborative); dedup ratio and bytes-per-edit (a regression means a chunking bug is quietly multiplying our storage and bandwidth bills); journal/notifier lag and 409 rate per namespace (hot-partition early warning). Rollouts: the sync client ships staged (1% → 10% → all) with the protocol versioned, because the scariest deploy in this company is a client bug that mass-produces wrong manifests — server-side validation and the version-history escape hatch are the blast-radius limiters.
How interviewers actually probe this design once it's on the board:
| Decision | Why | What it costs |
|---|---|---|
| Metadata plane / data plane split; presigned direct block upload | Two workloads with opposite needs; bytes never transit app servers; data plane swappable (Magic Pocket) | Two systems to operate; client is smarter and therefore buggier |
| Fixed 4 MB blocks over content-defined chunking | Predictable offsets, cheap CPU, covers in-place/append edits — most real edits | Insert-heavy files re-upload from the edit point onward; CDC held as a measured upgrade |
| Client-visible dedup scoped per account; cross-account dedup only inside the storage layer | Hash never becomes a bearer token (Dropship); no "does anyone store this file?" oracle; storage savings fully retained | First upload of a popular file pays full bandwidth per account |
| Base-version CAS at commit; conflicted copy on failure; no LWW, no auto-merge | Detects concurrency exactly; preserves both humans' work; matches the durability promise | Users see ugly duplicate files and must merge by hand |
| Metadata in sharded SQL, sharded by namespace | Commit CAS is a single-shard transaction; shared folders stay co-located | Giant shared namespaces become hot shards needing cache/batch mitigation |
| Refcount GC with leases + grace window | Closes the check/GC race; also sweeps orphan blocks | Dead bytes linger for the grace period — paid storage for safety |
A junior asks: "Why does Dropbox make those annoying 'conflicted copy' files instead of just merging, and how does it even know a conflict happened?" Explain in five or six sentences.
Every edit you sync carries a note saying which version it was based on, and the server only accepts the edit if that's still the newest version — like "I'm updating draft 7" being rejected because someone already made draft 8. When that rejection happens, the server has proof that two people edited the same starting point without seeing each other's work: a real conflict, not a guess. Now, most files are things like .docx or .xlsx — binary formats where smashing two versions together byte-by-byte produces a corrupted file, so automatic merging isn't cautious, it's impossible to do safely. Picking a winner by timestamp is worse: whoever synced last would silently erase the other person's work, and they'd only find out when the paragraph they wrote is just gone. So Dropbox does the only honest thing — it keeps both versions and names the second one "conflicted copy" with the device and date, so a human can merge with full information. The file is annoying precisely because it's evidence the system refused to destroy someone's work quietly.
Now add: camera-roll backup from phones — millions of users, each with tens of thousands of small photos, on batteries and metered connections. What in the design strains, and what do you change?
The strain is metadata, not bytes: a 3 MB photo is one block, so blob storage is easy, but 30,000 photos = 30,000 file rows, manifests, and journal entries — per-file commits would hammer the metadata shard and drain the battery with radio wake-ups. Changes: batch commits (one metadata transaction covering hundreds of new files), delta sync stops mattering (photos never get edited in place — write-once), dedup stays valuable (the same photo shared through five apps), and the client scheduler becomes a real component: upload on Wi-Fi + charging by default. Small-file-heavy workloads can also pack many tiny files' bytes into shared container blocks to keep blob-store object counts sane. The core insight: this workload flips the system from bandwidth-bound to metadata-QPS-bound, so the optimizations move planes.
10x the file size: a video-production team syncs 500 GB project files. Walk through where the current design creaks, number by number.
500 GB / 4 MB = 125,000 blocks. The check call now carries 125K hashes (~4 MB of hashes — page it); the manifest is similarly huge (store it chunked, or as a Merkle tree so two versions can be diffed without reading 125K entries). Upload time at 100 Mbps is ~11 hours, so the upload session and dedup leases must survive sleeps and IP changes for days — session TTLs and the GC grace window must stretch accordingly. The commit CAS itself doesn't care about file size — that's the payoff of manifest-last. Consider larger blocks (16–64 MB) for huge files to cut per-block overhead 4–16x, at the cost of coarser dedup and delta granularity — a per-file-size-tier block size is a reasonable committed answer.
Requirement flip: the PM wants Google-Docs-style live co-editing inside this system. What survives from your design and what must be rebuilt?
Almost nothing on the write path survives, and saying so is the right answer. This design syncs opaque files with whole-version granularity: conflicts are detected at commit and resolved by forking, on a timescale of seconds to days. Live co-editing needs operation-granularity sync — keystrokes as operational transforms or CRDT updates (p2c04) — flowing through a server that understands the document's structure, with conflicts merged in milliseconds, automatically, because the ops are designed to commute. The blob store, sharing/ACL model, and version-snapshot machinery survive as the persistence layer underneath. This is why Dropbox built Paper as a separate product rather than a feature of file sync: same company, different machine, because "converge two offline binaries" and "merge concurrent keystrokes" are different problems that only sound alike.
Requirement flip: a new enterprise tier demands end-to-end encryption — the server may never see plaintext. Which parts of your design die, and what's your least-bad rescue?
Cross-account dedup dies cleanly: encrypting the same block under different user keys yields different ciphertexts, so identical files stop deduplicating across accounts — the storage economics from the estimation take the hit, and you should quote it. Per-account dedup and delta sync can survive: the client encrypts per-block with an account key, and since the client sees plaintext it still computes hashes and diffs before encrypting. The tempting rescue — convergent encryption, deriving the key from the content hash so identical plaintexts encrypt identically and dedup revives — reintroduces the Dropship-style oracle: anyone can test whether a specific known file exists in the store. Least-bad committed answer: E2E tier = per-account dedup only, priced accordingly; the conflict machinery is untouched because CAS and manifests never needed plaintext.
Estimation drill: your notifier tier holds one long-poll connection per active device. At 100M daily-active devices with a 60-second poll cycle, size the tier — and name the failure mode that dwarfs steady state.
Steady state is mild: ~100M held connections (mostly idle — a tuned server holds ~1M mostly-idle connections, so ~100–150 machines with headroom) and 100M/60 ≈ 1.7M polls/sec of trivial "anything new?" checks. The killer is synchronized reconnection: a notifier-tier deploy, a regional network blip, or a mobile-carrier hiccup disconnects tens of millions of devices that all reconnect at once — a self-inflicted thundering herd 100x steady state, which also stampedes the changes endpoint behind it. Defenses: client reconnect jitter (randomized backoff spreading reconnects over minutes), connection draining on deploys, and admission control at the edge (p2c07). The senior tell is naming the herd unprompted — steady-state math on this tier is the easy 20%.
Friday, 10:00am. Coldplay announces one stadium show. Tickets go live at noon. By 11:55, ten million people are hammering refresh. At 12:00:00, the seat map opens — and somewhere in that crowd, four thousand people tap the same seat, 14B, within the same second. Exactly one of them can have it.
Now the interview version. The interviewer says: "Design Ticketmaster." Most candidates hear "huge traffic" and reach for the scaling toolbox — autoscaling groups, sharding, CDNs. Twenty minutes later they've built a system that can serve the seat map to a hundred million people and still sells seat 14B twice. The interviewer was never worried about traffic. Traffic is the decoy.
Selling one seat twice isn't a latency blip. It's two humans at a stadium gate holding tickets for the same chair — refunds, press coverage, and in one real case we'll get to, a national consumer-protection agency ruling that your system oversold a stadium. This walkthrough is built around the deep dive that decides your level: letting thousands race for the same database row and guaranteeing exactly one winner.
I'd scope this out loud in the first two minutes. Functional requirements, committed:
Out of scope, said explicitly: resale marketplace, dynamic pricing, venue entry scanning, search. Say “Fair?”, get the nod, and only then draw a box.
Non-functional requirements, with numbers I'll defend:
The numbers here don't size a fleet — they change the shape of the machine. Watch what falls out.
10 million buyers, 50,000 seats. That's a 200:1 ratio: 99.5% of my "traffic" cannot result in a sale no matter how well the system works. Scaling the purchase path to 10M concurrent users means scaling to process 9.95M guaranteed failures. The number screams a conclusion: don't autoscale the store — control admission to it. Trickle users into the buying flow; let everyone else wait in a queue that costs almost nothing to hold.
Write volume is tiny. A 30-minute sellout is 50K sold seats in 1,800 seconds — about 28 purchases per second. Add expiring and retried holds, call it a few hundred inventory writes per second. One decent Postgres does this without breathing hard; sharding the inventory solves a problem I don't have. The fire isn't write throughput — it's write contention: thousands of those writes aim at the same few rows in the same second.
The seat map is small. 50K seats × ~100 bytes of state ≈ 5 MB. The entire inventory of the hottest event on Earth fits in one cache node's spare change. But if 10M waiting users each poll the map every 5 seconds, that's 2M reads/sec — so the map must be served from cache/CDN as a slightly stale snapshot, and only admitted users get the fresher view.
Three numbers, three decisions: admission control instead of autoscaling, one strong database instead of a sharded fleet, cached-stale reads instead of live ones. That's estimation doing its job.
GET /events/{id}
GET /events/{id}/seatmap → cached snapshot, seconds stale
POST /events/{id}/holds {seat_ids: [...]} → hold_id, expires_at
POST /holds/{hold_id}/purchase Idempotency-Key: <uuid> → order
DELETE /holds/{hold_id} (user changed their mind)
Core tables. events and orders are routine. The interesting one is seats — one row per physical seat per event:
seats(event_id, seat_id,
status, -- free | held | sold
holder_user_id,
hold_expires_at,
order_id)
Every hard problem ahead is a fight over who gets to flip status on one of these rows. The hold endpoint returns an explicit expires_at so the client can show a countdown — honest UX comes free when the data model is honest.
Read path: a user (queued or admitted) fetches the seat map from cache. The snapshot is rebuilt from the inventory DB every couple of seconds and shows held seats as a best-effort overlay. It is deliberately, honestly stale — more on why that's fine below.
Write path: an admitted user picks seats → POST /holds → one atomic conditional write flips the rows to held with a 10-minute expiry → user pays → success flips held → sold, writes the order, and drops a message on the ticket queue for issuance and email. Payment failure releases the hold. The loop is closed: thumb to ticket, and every arrow has a failure story we'll cover.
The consistency commitment, stated as I'd state it at the whiteboard: the inventory database is the single source of truth and it is CP — one Postgres primary with synchronous replica failover; if it's briefly unavailable, purchases fail loudly and safely. Everything else — event pages, seat maps, queue positions — is AP: cached, seconds stale, because a wrong answer there costs a click, not a double-sold seat.
Why is a stale seat map acceptable? Because even a perfectly live map is stale by the time you act on it. The map says 14B is free; 300 milliseconds later, when your tap arrives, someone else took it. The map was always a hint; the conditional write is the truth. The "seat taken, pick another" experience exists at every staleness level, so I'll pay a few seconds of staleness to take 2M reads/sec off the database doing the one job that must not fail.
The seat map is the "seats available" sign outside a cinema hall; the inventory DB is the clerk at the counter. The sign might lag by a minute, and nobody considers that lying — you still have to ask the clerk, and the clerk's answer is final. Trouble only starts if you let people buy tickets from the sign. Where the analogy breaks: our clerk must survive four thousand people asking for the same seat in one second, which no human counter has ever faced.
Thousands of requests will try to buy the same seat in the same second, and exactly one may win. This is not a traffic problem — it's a concurrency-correctness problem on a handful of burning-hot database rows. Every naive design has a read-then-write gap where two buyers both see "free" and both pay. The interviewer invented this question to watch you find that gap, close it with an atomic operation, and then handle the ugly edge it creates: what happens when a hold expires while its owner's payment is mid-flight. Spend your deep-dive time here.
Walk the naive version concretely, because the bug hides in code that looks obviously correct:
row = SELECT status FROM seats WHERE seat_id='14B' -- says 'free'
if row.status == 'free':
UPDATE seats SET status='sold', ... -- so we sell it
Request A reads at 12:00:00.100 — free. Request B reads at 12:00:00.103 — also free, because A hasn't written yet. A writes sold. B writes sold. Both commit. Both users get charged, both get confirmation emails, and your invariant is dead. Under read-committed isolation — the default in Postgres, the one your ORM gives you — nothing stops this. The check and the write are two separate operations, and the whole race lives in the gap between them. This is the p1c03 isolation lesson wearing a concert t-shirt.
Rung 1: pessimistic locking. SELECT ... FOR UPDATE takes a row lock while reading, so B's read blocks until A's transaction finishes, then sees sold and gives up. Correct — and a fine answer for a normal booking site. But feel the drop-day mechanics: four thousand transactions queue on 14B's row lock, each waiter pinning a database connection, and the pool has maybe 200. Lock queues back up into connection starvation, and now requests for other seats can't get a connection either. Pessimistic locking has a contention ceiling, and this problem lives above it.
Rung 2: optimistic concurrency. Add a version column; write with UPDATE ... WHERE version = 7; zero rows matched means someone beat you — re-read and retry. Beautiful at low contention, and the right tool elsewhere in this book. Here it's actively wrong, and saying why is the senior move: optimistic concurrency assumes conflicts are rare, and ours are the whole point. 4,000 contenders race; 1 wins; 3,999 retry against a seat that's already gone — or onto the next hot seat, and lose again. One spike becomes a retry storm several times its size, aimed at your hottest rows. Optimistic locking under maximal contention is a load amplifier.
Rung 3 — my commitment: reservation with TTL via one atomic conditional write. Don't read then write. Make the claim itself the atomic operation:
UPDATE seats
SET status='held', holder_user_id=:u,
hold_expires_at = now() + interval '10 minutes'
WHERE event_id=:e AND seat_id='14B'
AND (status='free'
OR (status='held' AND hold_expires_at < now()))
One statement. The database checks the condition and applies the write as a single atomic action — no gap for a rival to slip through. Rows affected = 1: the seat is yours, come pay. Rows affected = 0: seat's gone, fail instantly — no retry storm, losers cost one cheap statement. The same trick works as a Redis SET NX PX or a DynamoDB conditional put; I'm committing to Postgres because the money path deserves real durability and transactions, and our write volume is a rounding error for it.
Notice the second clause: an expired hold counts as free. That's lazy expiry — the next claimant reclaims a dead hold in the same atomic write, no janitor required for correctness. I'd still run a background sweeper flipping long-expired holds back to free so the cached seat map doesn't show ghosts, but the sweeper is cosmetic; the conditional write is the law.
A hold is a fitting room. You take the shirt in, and for ten minutes it is off the rack — nobody else can grab it, and the store didn't have to decide whether you'll actually buy it. Dawdle past your time and staff quietly put it back on the rack for the next person. The one rule that makes the whole shop work: taking the shirt off the rack must be a single grab. If checking-the-rack and taking-the-shirt are two separate motions, two shoppers end up holding one sleeve each.
Now the timeline the interviewer is waiting for someone to notice. A user holds 14B at 12:00 and hits "pay" at 12:09:30. The provider is slow today. At 12:10:00 the hold expires — and lazy expiry means someone else can now legitimately claim 14B. At 12:10:40 the original payment succeeds. Two people have paid for one seat, and every component behaved exactly as designed.
I resolve this explicitly, in two layers, and I commit to both:
held → pending_payment (conditional write again: only if it's still their unexpired hold) and extend the expiry by the payment timeout plus margin — say 3 minutes. A pending_payment seat is not reclaimable. This closes the window for every payment that finishes inside the timeout, which is nearly all of them.On December 9, 2022, Bad Bunny played the 87,000-capacity Estadio Azteca in Mexico City — and thousands of fans holding tickets bought through Ticketmaster Mexico were turned away at the gates, their tickets scanned as invalid or duplicates. Ticketmaster's first explanation was counterfeits. Mexico's consumer-protection agency Profeco investigated and concluded otherwise: the tickets were genuine, and the system had oversold the venue. Profeco said it could fine Ticketmaster up to 10% of its annual sales in Mexico; Ticketmaster ended up refunding 2,155 affected customers 100% of the ticket price plus 20% compensation — about 18.2 million pesos, roughly a million US dollars. Notice the shape of the failure: not a crashed website, but duplicate valid-looking tickets for the same physical seats — the invariant this whole chapter defends. And the resolution matched our reconciliation rule: once overselling reaches humans, all you can do is refund with interest and eat the reputational bill. The database guard is cheaper.
Back to the 200:1 ratio. The inventory island and the store servers can comfortably serve a few thousand active shoppers. Ten million people are outside. The waiting room's job is to turn an uncontrolled stampede into a controlled drip, and to be fair and honest while doing it.
Mechanics I'd commit to: when the drop opens, arrivals land on a lightweight queue page — static assets from the CDN, one tiny API call to register presence. Admission into the store is a token bucket: the store publishes how many new shoppers per minute it can absorb (a dial, not a constant — say 2,000/min to start), and the queue service issues that many signed admission tokens per minute. The store's load balancer admits only requests bearing a valid token — signed (an HMAC, a keyed signature verifiable statelessly) so nobody can forge their way in, short-lived so tokens can't be stockpiled or resold.
Now the fairness question, a real design decision: who gets admitted first? Strict FIFO by arrival time sounds fair and isn't — it rewards whoever clicked fastest at 11:59:59.9, which in practice means bots and the best connections. I'll commit to the standard industry answer: everyone arriving within the opening window (say the first 5 minutes) is placed in randomized order — a lottery over the surge — and later arrivals join FIFO behind them. Then be honest in the UX: show real position and a real estimate ("~41,000 ahead of you, roughly 24 minutes"), never silently reshuffle. A queue people trust is one they'll wait in; one that looks rigged generates its own DDoS of rage-refreshing — and refreshing must never change your position, which the signed token also guarantees.
The waiting room is also the bot chokepoint. Rate-limit queue joins per account, per payment method, per IP, and per device fingerprint (a hash of browser characteristics — imperfect, but it raises the cost of pretending to be 500 users). For monster drops, add pre-registration: verify humans days before, with no time pressure, and admit only pre-verified accounts. That's Ticketmaster's Verified Fan program — which brings us to the day it wasn't enough.
On November 15, 2022, Ticketmaster opened the Verified Fan presale for the Eras Tour. The plan was textbook admission control: 3.5 million fans pre-registered, about 1.5 million got invite codes sized to what the system could handle. What showed up, by Ticketmaster's own account, was around 14 million — code-holders, code-less hopefuls, and a "staggering number of bot attacks" — generating a record 3.5 billion system requests that day, roughly four times their previous record. The site crashed repeatedly, sales were paused mid-drop, fans queued for hours, and days later Ticketmaster canceled the general public sale outright, citing insufficient remaining inventory. Taylor Swift called watching it "excruciating"; the US Senate held hearings. The architecture lesson is precise: the admission plan was sized to invited demand, but admission wasn't fully enforced at the edge — uninvited users and bots still reached the real infrastructure. A waiting room only works as the only door, standing in front of everything, with the cheap static queue page absorbing the millions who will never get in.
Purchase is a multi-step operation across two systems that can each fail independently: our inventory and an external payment provider. That's a saga — a sequence of local steps, each with a compensating undo — because I can't wrap Stripe in my database transaction (p2c01's whole point).
pending_payment (conditional write, expiry extended).pending_payment → sold, write the order, drop a message on the ticket queue.free, tell the user. The seat re-enters inventory within seconds, not after a 10-minute TTL funeral.The failure that matters isn't "card declined" — that's a clean no: compensate, done. It's the timeout: we asked the provider to charge and heard nothing. Did it happen? Unknown, and guessing in either direction is the error. The idempotency key makes it safe to ask again — retry the identical request or query status; the answer is deterministic. Until we know, the seat stays pending_payment. If the provider stays dark past our window, we fail the purchase, release the seat, and let reconciliation refund any charge that later turns out to have succeeded — same rule: refund money, never un-sell seats.
And the systemic version: the provider doesn't die, it gets slow — 30-second responses on drop day. Every in-flight purchase now pins a pending_payment seat for minutes. Pending seats pile up, the map fills with unbuyable seats, conversion craters — while the queue keeps admitting shoppers into an empty-shelved store. Defenses, in order: aggressive payment timeouts (fail at 10s, not 60s), a circuit breaker (p2c07) that stops feeding a drowning provider and fails fast, and — the underrated one — turn the admission dial down when conversion drops. The waiting room isn't just an entry gate; it's the system's backpressure valve. Admit fewer shoppers while payments are sick, and the queue absorbs the pain invisibly.
What I'd watch on a dashboard during a drop, unprompted:
sold seats against orders and payments: every sold seat has exactly one order, every order one successful charge, no seat twice. Expected result: zero, forever. Any nonzero pages a human immediately — the invariant's smoke detector, which also catches money-side ghosts (charges with no seat → auto-refund queue).Rollout: never let a real drop be the first test. Replay a synthetic drop — millions of simulated queue joiners, scripted contention on the same hot seats — against production infrastructure before every major on-sale. Ship changes behind flags, canaried on small-venue events (a Tuesday comedy club sale, not Coldplay). Admission starts conservative and ramps: you can always speed a queue up; over-admitting then slowing down means users watching checkouts time out, which is how trust dies.
What breaks at 10x? Say 100M users show up, same 50K seats. Here's the payoff of admission control: the store, inventory island, and payment path see exactly the same load as before — the token bucket doesn't care how long the line is. What must scale 10x is only the cheap part: the CDN-served queue page and the queue service (an append-only list plus a counter — Redis does this in its sleep, and it shards trivially because queue position doesn't need global precision). The part that's hard to scale was made load-independent; the part that must absorb arbitrary load was made trivially scalable. If the interviewer 10x's the seats instead, the seat-row model still holds — it's when seats become indistinguishable (general admission) that the model should change, and that's an exercise below.
The interviewer will let you talk traffic for a while, then aim every probe at the invariant and its edges:
| Decision | Chosen over | What it costs me |
|---|---|---|
| Atomic conditional write for holds | Row locks / optimistic retries | Losers get a hard "seat gone" instead of waiting for a maybe; every state change must be written as a guarded update, which is easy to get subtly wrong in review |
| Hold-with-TTL before payment | Direct buy (lock at purchase) | Inventory sits unavailable inside hold windows; abandoned carts delay resale of hot seats by up to 10 minutes |
| One CP inventory island (single Postgres) | Sharded / multi-leader inventory | A hard write-throughput ceiling and a real failover story to operate; fine at hundreds of writes/sec, revisit if an event ever needs 100x that |
| Stale cached seat map (seconds) | Live map per user | Users sometimes click seats that are already gone and see "seat taken" — a UX papercut accepted to shed ~2M reads/sec |
| Waiting room with randomized-window entry | Autoscaling the store; strict FIFO | People wait visibly, and randomization means arriving first in the surge guarantees nothing — honesty in the UI is the mitigation |
| Extend-on-payment-start + refund-late-success | Strict expiry (seat always reclaimable at TTL) | A slow payment can pin a seat a few extra minutes; rare double-payments become refunds and apology emails instead of double-sold seats |
A junior teammate asks: "We check if the seat is free before we sell it — how can it possibly sell twice?" Explain the race and the fix in five or six sentences.
Your check and your write are two separate trips to the database, and the world can change between them. If two requests both run the check while the seat is still free, both pass, then both write — the seat sells twice, and no default database setting stops it, because each request was individually valid. The fix is to stop checking first: one single statement that says "set this seat to held only if it is still free," so the database evaluates the condition and applies the write as one atomic action. Now there is no gap: the first statement wins, everyone else's matches zero rows and instantly knows it lost. The winner gets a 10-minute hold to pay, with an expiry so an abandoned cart returns the seat automatically. One subtlety: if a hold would expire while its owner's payment is processing, we extend it when payment starts — and if money and seats still disagree, we refund the money, because you can reverse a charge but not a seat someone else now owns.
New requirement: "best available" — the user asks for 3 adjacent seats and the system picks them. Your conditional write claims one row at a time. What goes wrong, and how do you extend the design?
Three independent conditional writes can partially succeed — you win 12C and 12D but lose 12E — leaving a broken pair you must release (a mini-saga) and retry elsewhere. Two clean fixes. One: wrap the three guarded updates in a single database transaction — all three match or roll back; easy in Postgres since every row lives in one place (a quiet payoff of the unsharded CP island). Two, cuter: precompute adjacency groups and claim a group row atomically, materializing the seat holds under it. Either way, retry alternative groups server-side, and mind fragmentation: greedy best-available allocation strands single seats between sold pairs — an inventory-yield problem venues genuinely care about.
Flip a requirement: the event is general admission — 50,000 identical tickets, no seats. Does the per-seat design survive? What would you change?
The per-row model collapses into one row: a counter at 50,000, and every buyer contends on it — the ultimate hot partition (p2c06). The conditional write still works (UPDATE ... SET remaining = remaining - 1 WHERE remaining > 0) but serializes all sales through one row's lock. Fix: split the counter into ~20 shards of 2,500; a buyer decrements a random shard, retrying another on zero — contention drops 20x, correctness holds because each shard is independently guarded. Alternatively, pre-generate 50K ticket IDs in a queue and pop one per admitted buyer — a pop is atomic by construction. Deeper lesson: with no scarce specific seat to protect while a human deliberates, holds matter less — you can collapse hold+buy into one step.
Your payment provider offers no status-query API and no webhooks — you send a charge, and if it times out you simply cannot ask what happened. Redesign the money path to stay safe.
You've lost the "ask again" escape hatch, so unknown outcomes must be resolved by structure. In order of preference: (1) refuse the constraint in real life — a provider without idempotent retry or status query is unfit for this system, and saying so is a legitimate senior answer; (2) two-phase money: authorize first, capture only after the seat flips to sold — an uncaptured authorization expires harmlessly, so an ambiguous timeout defaults to "customer not charged," the safe direction; (3) failing that, treat every timeout as failed, release the seat, and reconcile daily against the provider's settlement file, refunding any charge that snuck through. The ranking guiding all three: never oversell > never silently keep money > never make the customer retry unnecessarily.
Design the reconciliation job from the ops section concretely: what does it compare, how often does it run, and what does it do on each kind of mismatch?
Three datasets: seats (status + order_id), orders, and the provider's charge records. Invariants: every sold seat → exactly one order → exactly one successful charge; no seat in two orders; no charge without an order. Run a fast in-DB pass every few minutes during a drop (seats vs orders — one join over 50K rows) and a full three-way pass hourly plus end-of-day, since webhooks arrive late. On mismatch: seat sold twice → page immediately, freeze the pair, resolve before gates open — the never-event; charge without seat → auto-refund with apology, no human needed; order without charge → re-query via idempotency key, then mark paid or release. Log every auto-action to an append-only audit table — when a Profeco-shaped regulator comes asking, that trail is your defense.
"Design a web crawler." The interviewer says it casually, like they're asking for a for-loop. That's the trap. A crawler sounds like a for-loop: fetch a page, pull out the links, fetch those, repeat. Half the candidates who hear this question relax. Wrong move.
Here's what makes this problem different. When you design Instagram badly, your users suffer. When you design a crawler badly, other people's servers suffer. Your code knocks on the door of millions of machines you don't own — including a retired teacher's WordPress blog on a $5 VPS. Send that blog fifty requests a second and you didn't build a crawler; you built a small denial-of-service attack with a user agent string.
So this problem grades two things at once: can you run a pipeline at eleven thousand fetches per second, and can you make that pipeline incapable of being rude to any single host — in the same data structure. That structure is the URL frontier, and it's where this interview is won or lost. Everything else — storage, parsing, DNS — is supporting cast.
"A web crawler" is fuzzy on purpose. I'd scope it out loud: "I'll design a crawler that feeds a search index — the Googlebot/Common Crawl shape. Not a scraper for one site, not a JavaScript-rendering browser farm. Fair?"
Functional requirements:
Non-functional requirements — with numbers I'll defend:
robots.txt, honor per-host crawl delays, never overload any host. I'll say this now and repeat it later: every other property of this system can degrade under load; this one cannot.Fetch rate. 1B pages/day ÷ 86,400 seconds ≈ 11,500 fetches/second sustained. Web fetches are slow — 1–2 seconds average — so roughly 15–25K requests are in flight at any instant. That's an async-I/O problem, not a threads problem: one event-loop machine holds thousands of open connections, so fetching needs maybe a few dozen machines. Conclusion: fetch capacity is cheap. Whatever is hard here, it isn't "can we open enough sockets."
Bandwidth and storage. At ~100KB per HTML page, 1B pages/day is ~100TB/day ≈ 9–10Gbit/s sustained ingress. That forces fetchers spread across machines and uplinks, and compression before storage (HTML squeezes ~4x → ~25TB/day durable; petabytes per monthly cycle). Conclusion: content goes in a blob store, and only skinny metadata goes anywhere queryable.
The politeness math — the number most candidates never compute. Suppose the polite delay is 2 seconds per host. One host then yields at most 0.5 fetches/sec, so sustaining 11,500 fetches/sec needs at least 11,500 × 2 = 23,000 distinct hosts ready at every moment — realistically hundreds of thousands in rotation, since many hosts have only a few pages queued. This one calculation reshapes the design: throughput comes from breadth across hosts, never depth on one. A crawler isn't a firehose; it's a million polite trickles, and the frontier's job is keeping enough of them flowing.
One more hidden number: 11,500 fetches/sec means up to 11,500 hostname lookups/sec. Hold that thought — DNS gets its own section, because it has quietly wrecked real crawlers.
There's barely a public API — the "users" are internal systems. POST /seeds to inject starting URLs with a priority, and downstream consumers tail a "page fetched" stream (Part 2 streams pattern: the crawler produces, the indexer consumes, neither knows the other's pace).
Two data shapes matter. The URL record, keyed by a hash of the normalized URL: the URL, host, priority score, depth from seed, last-fetch time, ETag/Last-Modified, estimated change frequency, failure count. The page store: content in a blob store keyed by content hash (identical bodies stored once), plus a metadata record mapping URL → content hash, fetch time, HTTP status, outlinks. Common Crawl — the nonprofit publishing an open crawl of roughly 2 billion pages every month — stores exactly this shape as append-only WARC archives: the raw HTTP response plus capture metadata, concatenated into big files. Append-only is right for us too: a crawl is a log of observations, not a table you update in place.
Walk the loop once. Seeds enter the frontier — the giant prioritized, politeness-aware to-do list we'll dissect next. Fetcher workers ask it "what may I fetch right now?", check the cached robots.txt for that host, resolve DNS from a local cache, and download. Content gets checked for duplicates, written to the blob store, and a "page fetched" event goes onto a stream. Parser workers — a separate fleet, because parsing is CPU-bound while fetching is I/O-bound, so they scale independently — consume the stream, extract every <a href>, normalize the URLs, and push them through the URL-seen filter. Genuinely new URLs get a priority score and enter the frontier. The loop closes: a crawl is this loop run a billion times a day without ever being rude.
Write path and read path, explicitly: the write path is the loop itself (frontier → fetch → store → parse → frontier). The read path is downstream consumers reading the blob store and the fetched-page stream — deliberately decoupled through the stream so a slow indexer can never backpressure the crawl.
The frontier must satisfy two masters that want opposite things. Priority says: fetch important, fresh-changing pages first — a new BBC article beats page 94,000 of a forum archive. Politeness says: however important a host's pages are, never hit it more than once per delay window. A naive priority queue fails politeness catastrophically: the top thousand entries usually belong to the same few important hosts, so fetchers drain the queue straight into one server's face. Naive round-robin over hosts fails priority: junk fetched at the same rate as gold. The interview IS the reconciliation of these two, at a scale — billions of queued URLs — where the structure must live on disk. Nail this section and the rest is a formality.
The classic answer — it comes from Mercator, a research crawler from Compaq's labs whose war story is coming up, and it still underlies serious designs — is a two-level queue.
Front queues: priority. A few dozen FIFO queues, one per priority band. The prioritizer scores each incoming URL — site importance (PageRank-ish, or simpler: how many distinct domains link in), expected change frequency, depth from seed — and drops it into the matching band. A selector pulls with biased randomness: high bands far more often, low bands never fully starved.
Back queues: politeness. A large pool of FIFO queues — several times more than fetcher workers — with the invariant everything depends on: each back queue holds URLs from exactly one host, enforced by a host→queue mapping table. A host lives in one queue, a queue is consumed sequentially, so requests to any host are serialized by construction. Politeness isn't a rate-limiter bolted on top that might have bugs; it's a structural property of the data layout. That's the sentence to say at the whiteboard.
The dance between the levels. When a back queue runs empty, it refills from the front queues: the selector pops a URL (biased toward hot bands) and checks the host map. If that host already owns another back queue, the URL is appended there and the selector draws again, until it finds an unmapped host — which the empty queue claims, updating the map. Net effect: important hosts stay continuously queued, and refills keep pulling fresh hosts into rotation — exactly the thousands of distinct ready hosts our politeness math demanded.
The timer heap: deciding "now." One structure remains: when may each host next be contacted? Keep a min-heap of (next_allowed_time, back_queue_id). A worker pops the root — the host whose wait expires soonest — waits out the last few milliseconds, fetches one URL from that queue, and re-inserts the entry with next_allowed_time = completion + delay. A queue's entry exists at one heap position, so a host can never be fetched concurrently, even by ten thousand workers.
What's the delay? I'd commit to max(robots crawl-delay, ~1s base, adaptive term). The adaptive term is the senior detail — a static delay treats a small VPS like a CDN. Scale it to the host's observed response time (a server answering in 3 seconds is straining; multiply its delay), and on 429/503 back off exponentially and honor Retry-After. A struggling server told you it's struggling; a polite crawler listens.
The frontier is a call center with a courtesy rule. You have a ranked list of people to reach (front queues — VIPs first), but you may never ring the same household more than once every two minutes. So you keep one call sheet per household (back queues) and a diary of "may call again at…" times (the timer heap). Agents always dial whoever's diary time comes up next; a finished sheet gets replaced by the next name off the ranked list. Priority decides who gets on a sheet; the diary decides when anyone gets dialed. The analogy breaks on scale: we need hundreds of thousands of "households," which is why the diary is a heap and the call sheets live on disk.
Scale reality check. Billions of queued URLs at ~80 bytes each is terabytes — the frontier is a disk structure with only the hot edges (queue heads, heap, host map) in memory. One machine can't hold it, which brings up the elegant bit of the whole design: shard the frontier by hash of hostname. Parsers publish extracted URLs onto a stream partitioned by host hash (straight out of Part 2), so every URL for a host routes to the one frontier node owning that host. Each node's back queues, timer heap, robots cache, and DNS cache concern only its own hosts — so the property that must never break, per-host delay, never requires distributed coordination. Compare the alternative someone always proposes: a central rate-limiter checked before every fetch — a network hop on the hottest path and a single point of failure guarding your only unbendable invariant. Ownership beats coordination; say it in exactly those words.
Before any fetch to a host, the crawler needs that host's robots.txt: disallowed paths, plus any Crawl-delay. Fetch it once per host and cache with a ~1-day TTL — never per-page (doubles fetch volume), never forever (owners change their minds, and crawling a newly disallowed path is how you end up in an angry blog post). On failure: 404 means no rules, crawl freely; 5xx or timeout means the site is having a bad day — "come back later," not "come on in." And identify honestly: a real User-Agent with a URL explaining who you are and how to reach you, as CCBot and Googlebot do — otherwise operators can't tell you from an attack, and will treat you like one.
Why so absolute? Two reasons, and I'd give both in the interview. Ethics: an impolite crawler is a DDoS with paperwork — 11,500 requests/second spread politely is invisible to everyone; pointed at one small host it's an outage you caused. Pragmatics: impolite crawlers get IP-banned, fed garbage, or blocked at the CDN level across millions of sites at once, and a banned crawler sees a distorted web that quietly poisons everything downstream. Politeness isn't a tax on throughput; it's the license to operate.
In July 2024, Read the Docs — host of documentation for a huge share of open-source projects — published "AI crawlers need to be more respectful," with receipts. One AI company's crawler had downloaded 73TB of zipped HTML from them in a single month, almost 10TB in one day, hammering the same large files hundreds of times from many IPs with no rate limiting and no conditional requests. The bandwidth bill: over $5,000, landed on a nonprofit. The crawler had no politeness layer at all — no per-host delay, no dedup of what it had already fetched, no backoff. Read the Docs traced it, emailed the company, got reimbursed — and joined the wave of sites blocking careless bots outright. This is the failure mode your frontier exists to make structurally impossible, and why "I'd rate-limit somewhere" fails at the whiteboard: the per-host delay must be enforced by the data structure, not by good intentions.
Dedup happens twice, for two different questions, with two different tools.
URL-seen: have I queued this URL before? Every page yields dozens of outlinks — tens of billions of candidates a day, the vast majority already known. First, normalize: lowercase the host, strip fragments and default ports, resolve relative paths, drop tracking parameters, sort the query string. Skip this and example.com/a?x=1&y=2 and EXAMPLE.com/a?y=2&x=1 count as two pages, and your "dedup" leaks duplicates all day. Then test membership in a set holding ~20 billion entries. Exact hashes would be hundreds of GB — awkward for RAM — so here's the payoff of Part 2's sketches chapter: a Bloom filter in front, a durable exact store behind. At ~10 bits per entry, 20B URLs fit in ~25GB of RAM (split across shards), answering "definitely new" in nanoseconds with ~1% false positives. The exact set — disk-based, updated in big sorted batches, sequential I/O only — is the source of truth for rebuilding the filter and for recrawl scheduling.
Now walk the false-positive consequence, because the interviewer will: a false positive means the filter says "seen" for a genuinely new URL, and we skip a page we've never crawled. Acceptable? For a search crawler, yes — I'd commit to it. We miss ~1% of new URLs on first sight; popular pages are linked from many places and get re-tested on every appearance. Note the asymmetry that makes this safe: Bloom filters never produce false negatives, so re-queuing a known URL — the expensive direction — is impossible. But flip the requirement to legal archiving with a completeness guarantee and the trade is no longer yours to make: "maybe seen" must then be verified against the exact store before dropping. Same architecture, one dial flipped — knowing which setting your requirements demand is the senior move.
Content-seen: have I fetched this exact content under a different URL? Mirrors, www vs bare domain, print views, session-ID spinners — a huge slice of fetched bytes is content you already have. Hash every body (SHA-256); the blob store is keyed by that hash, so an exact duplicate costs a metadata row and zero storage. Near-duplicates — same article, different sidebar — need fuzzier eyes: simhash, a 64-bit fingerprint built so similar documents differ in only a few bits. New term, so one plain sentence: where SHA-256 scrambles completely on a one-byte change, simhash makes small content changes cause small fingerprint changes, turning "is this 95% identical to something I have?" into a cheap bit comparison. Google published exactly this (Manku's 2007 paper): simhash over 8 billion pages, Hamming distance ≤ 3 = near-dup. Near-dups still get stored, but demoted in recrawl priority and flagged for the indexer.
Assume the web is trying to trap you, because parts of it are. The classics: a calendar widget whose "next month" link works forever — the crawler marches politely toward the year 3000. Session IDs in URLs, minting infinite "new" URLs for one page. Redirect loops. And spam farms: millions of auto-generated pages densely linking to each other, each one technically unique.
No single defense wins; you layer cheap ones. Depth caps: drop URLs more than ~15–20 hops from a seed or with absurd path lengths. Redirect caps: five hops, then abandon. Pattern heuristics: repeated path segments (/a/b/a/b/a/b), date paths marching past today, exploding query-parameter counts. Content-seen catches spinners — a thousand URLs, one content hash. But the defense that actually contains the damage is the per-domain budget: no domain consumes more than its earned share of the crawl, where the share scales with reputation — cheaply approximated by how many other domains link to it. A trap can generate infinite URLs, but it can't force you to spend fetches on them. Budgets convert "adversary controls my frontier" into "adversary wastes their own quota."
Budgets are how a tourist survives a city with infinite alleys. You don't photograph every brick in the first interesting alley — you allocate: an afternoon for the famous cathedral, an hour for the quirky side street, five minutes for the shop that's clearly a tourist trap. If a street performer keeps unveiling "one more act," your budget walks you away no matter how the acts multiply. The crawl budget does the same: attention proportional to earned importance, and no corner of the web — however infinitely it generates novelty — can consume more than its allocation.
In 2007, researchers at Texas A&M ran IRLbot — a crawler on a single server — for 41 days straight: 6.3 billion pages fetched, averaging ~1,789 pages/second, published at WWW 2008. Two findings shaped how everyone builds crawlers now. First, what broke crawlers at scale wasn't bandwidth but the URL-seen check: billions of membership tests overwhelmed RAM and naive disk lookups alike. Their fix — DRUM, for Disk Repository with Update Management — batched URL checks into large sequential disk merges instead of random reads — the ancestor of the batch-merge exact store behind our Bloom filter. Second, naive breadth-first crawling gets captured by spam: auto-generated link farms with massive branching factors flooded the frontier until real sites starved. Their answer — crawl budgets allocated by domain reputation, measured by how many other domains link in — is the direct ancestor of the budget defense above. One machine, 6.3 billion pages, and both of this chapter's hard problems documented in one paper.
Every fetch starts with resolving a hostname. At 11,500 fetches/second, forwarding every lookup upstream melts your resolver — and each remote resolution costs 20–200ms. The fix is boring and mandatory: a local caching resolver on every fetcher node, async lookups so slow resolutions never block the event loop, and pre-resolution while a host waits in the timer heap. Host-hash sharding helps here too — a host's lookups always hit the same node's cache, so hit rates are excellent.
Mercator was the research crawler built at Compaq's Systems Research Center by Allan Heydon and Marc Najork. Their 1999 paper's most quoted finding wasn't about queues at all: profiling showed DNS resolution consuming 87% of each thread's elapsed time. The culprit was a synchronized resolver interface — Java's, and gethostbyname underneath it, allowed one uncached lookup outstanding at a time — so their fully parallel crawler was quietly single-file through DNS. They wrote their own multi-threaded resolver firing parallel requests at a local nameserver, and DNS dropped from 87% to 25% of elapsed time. Their 2001 follow-up report is where the two-level frontier this chapter draws comes from, timer heap and all — right down to running about three times as many back queues as crawler threads, the same rule of thumb we used above. The DNS lesson generalizes: your architecture is only as parallel as its most serialized dependency, and that dependency is usually something so mundane you never drew it on the whiteboard.
The web doesn't hold still — a news homepage changes hourly; a 2009 blog post never will. Re-crawling everything at one rate wastes most of your billion daily fetches on unchanged pages. So the metadata store tracks each URL's change history, and the recrawl scheduler feeds URLs back into the front queues with priority ∝ importance × expected staleness: pages that changed on recent visits get visited more; pages that never change decay toward monthly. Make refreshes cheap with If-Modified-Since/If-None-Match — a 304 Not Modified costs a round trip instead of 100KB. Sitemaps, where published, are free intelligence: change frequencies and last-modified dates straight from the source.
A fetcher crash loses nothing. Fetchers are stateless; the frontier is the durable record (append-only queue files, checkpointed heap/map state). Workers take URLs on a lease; crash, and the lease expires and the URLs are dispensed again. At-least-once delivery, safe because a duplicate fetch is idempotent — content-hashed storage absorbs it. Exactly-once would cost coordination on the hottest path and buy nothing: re-downloading a page is wasted pennies, not corruption. Cheap crash-safety through idempotency — Part 2's pattern, and here it's the whole story.
A frontier shard dies. Its hosts go uncrawled until the shard recovers or a replacement replays its state — acceptable, because crawling is latency-tolerant; nothing user-facing is waiting. On takeover, politeness state restarts conservatively: assume every host was just fetched and owes a full delay. Err impolite-never, idle-briefly-fine.
What I'd monitor, in priority order:
Rollouts: new prioritizer or politeness logic canaries on a low-stakes shard, with the 429 rate as automatic rollback trigger. The kill switch is the frontier pausing dispensing: fetchers drain in-flight work and idle, the stream goes quiet, nothing breaks. Backpressure-friendly by construction.
What breaks at 10x? 10B pages/day = 115K fetches/sec and ~1PB/day of ingress. Fetchers and frontier shards scale linearly — the reward for zero cross-shard coordination. Three things don't. Bandwidth: 100Gbit/s sustained means multi-region fetch clusters, each owning a host-hash range (ownership stays exclusive — the invariant survives geography). The URL-seen store: 10x churn on the batch-merge pipeline needs re-tiering before it needs disks. And the subtle one: the politeness math now demands ~10x more distinct ready hosts at every instant — but host importance is zipfian — a few thousand hosts hold most of the pages anyone wants, and the tail falls off a cliff (Part 2's hot-partition lesson in a trench coat) — so past a point the constraint isn't your fleet, it's the web: the pages you most want sit on hosts whose fetch rate politeness caps. At 10x, the frontier's job shifts from "keep up" to "choose well."
Interviewers use the crawler to test whether you can serve two masters in one data structure, and whether you think about systems that touch strangers' machines. Expect these, nearly verbatim:
| Decision | Chose | Over | Cost I accept |
|---|---|---|---|
| Frontier structure | Two-level queues + per-host timer heap, disk-backed | Single global priority queue | More moving parts; refill logic is subtle and needs real tests |
| Politeness enforcement | Structural: one host ↔ one back queue ↔ one timer entry | Central rate-limiter service | Politeness logic duplicated per shard; cross-shard config drift possible |
| Distribution | Shard everything by host hash; zero cross-shard coordination | Shared state / coordination service | A dead shard's hosts pause until recovery; hot shards need occasional range rebalancing |
| URL-seen | Bloom front + batch-merged exact store | Exact check per URL | ~1% of new URLs skipped at first sight; unacceptable if completeness is contractual |
| Fetch semantics | At-least-once + idempotent content-hash store | Exactly-once | Occasional duplicate fetch (wasted bandwidth, slight politeness cost) |
| Content storage | Append-only blobs keyed by content hash | Mutable per-URL rows | Compaction/GC of superseded captures becomes a background job |
Explain the URL frontier to a junior in five or six sentences — specifically why a plain priority queue isn't enough and how the two-level design fixes it.
A crawler keeps a giant to-do list of URLs, and two rules fight over it: fetch important pages first, but never hit the same website too often. A plain priority queue obeys only the first rule — the top of the queue is usually many URLs from the same important site, so your workers would hammer that one server. The fix is two levels: front queues sort URLs by priority, and back queues hold URLs for exactly one website each, so requests to any site line up single-file by construction. A little clock (a min-heap of "may contact again at…" times) tells workers which site's turn it is, and after each fetch the site goes back in the clock with a fresh delay. When a back queue empties, it refills from the front queues — that's where priority sneaks back in: important sites get back-queue slots more often. Priority decides what enters the rotation; the per-site queues and clock decide when anything is actually fetched.
Politeness math drill: your polite delay is 2 seconds per host. (a) How many distinct hosts must be ready at every instant to sustain 11,500 fetches/sec? (b) A single host has 10 million pages — how long does a full crawl of it take, and what does that force in your design?
(a) Each host yields at most 0.5 fetches/sec, so at least 11,500 × 2 = 23,000 hosts mid-rotation — in practice far more, since many back queues hold few URLs and spend time refilling. (b) 10M × 2s = 20M seconds ≈ 231 days for one host. "Crawl everything on big hosts" is arithmetic fiction: budget the host, prioritize within it (homepage and hubs before page 9 million), and use sitemaps plus conditional GETs to spend its tiny allowance only on what changed. The delay isn't a throughput tax — it's a cap on per-host coverage, a much deeper constraint.
Now add JavaScript rendering: half the modern web builds its content client-side, and your indexer wants what users actually see. What changes in your design?
Rendering means a headless browser per page — roughly 100x the CPU and memory of a plain fetch, and each render fires dozens of sub-requests (JS, CSS, APIs) that must obey politeness for their hosts too. Unaffordable for a billion pages a day, so: plain fetch first, a cheap classifier flags JS-dependent pages (near-empty body, framework markers), and only those enter a render queue with its own budget — rendering is a privilege important pages earn, not a default. The frontier is unchanged; you've added a more expensive fetcher class behind the same politeness machinery. This mirrors Googlebot: crawling and rendering are separate phases, with rendering fed by its own queue at its own pace — sometimes seconds behind the fetch, sometimes much longer, but never the same budget.
A site owner emails: "your crawler is hitting my server 30 times a second" — but your logs show every one of their hostnames correctly honoring its 2-second delay. What happened, and what's the fix?
Politeness per hostname isn't politeness per server. The owner runs 60 hostnames — customer subdomains, country domains — on one backend, and 60 polite queues at 0.5 req/s each add up to 30 req/s at the shared origin. Fix: a second politeness layer keyed by resolved IP, so hosts sharing an origin share a rate budget — your DNS cache already has the data. The subtle cost: CDNs put thousands of unrelated sites behind shared IPs, so the IP-level cap must be much looser, with the strict per-hostname delay as the primary rule. The question tests whether you know what politeness is for: protecting machines, not name strings.
Requirement flip: instead of a discovery crawler for search, you now need to re-check 100K known product pages every 10 minutes for price changes. Which parts of your design survive, and which collapse?
Almost everything shrinks except politeness. 100K pages / 600s ≈ 170 fetches/sec — a couple of machines. No discovery means no link-extraction loop, no URL-seen problem, no Bloom filter, no traps or budgets: the frontier degenerates into a recurring schedule — a cron wheel. What survives untouched: the per-host timer heap and delays (100K product pages span a few thousand retailer hosts, and commercial sites block scrapers fastest of all), robots.txt handling, conditional GETs, and content hashing (only act on changes). Say the lesson out loud: politeness machinery is the invariant core of every crawler; the frontier's priority half is what morphs with the use case.
"Design a rate limiter for our public API." The candidate before you smiled at this question. A counter, an if statement, a 429 — the HTTP status that means "you're over your limit, back off." Ten minutes, done. That candidate did not get the senior grade, because this question is a trap with a friendly face. The counter is trivial. What's hard is where the counter lives.
Here's the moment the question is built around. Your API runs behind 40 gateways spread across three continents. A customer's key is limited to 100 requests per second. That customer's traffic — or an abuser holding their stolen key — arrives through anycast routing — one IP address advertised from every location at once, so each request gets pulled to whichever gateway is nearest to it — which means their requests land on all 40 gateways. Each gateway sees a gentle trickle: two or three requests per second. Every local counter says "fine." Globally, the key is doing 4,000 requests per second, and your database is on fire.
So the real question is: how do 40 machines on three continents agree on one number, thousands of times per second, without adding latency to every single API call your company serves? That's distributed counting under a latency budget of about one millisecond. It has no perfect answer — which is exactly why interviewers love it. The senior move is to name the impossible corner early and negotiate what you'll sacrifice.
Functional requirements I'd state and get agreement on:
Non-functional, with numbers I'll commit to:
Most estimation in this problem exists to kill one idea: the single central counter.
Throughput. The limiter runs on every request, so it does 1M checks/sec at peak. Whatever the check costs, multiply by a million per second, forever. This is why the check must be nearly free.
Latency. A counter in one central place means every request pays a round trip to it. Cross-continent RTT is 80–150ms — my budget is 1ms, so a global synchronous counter is dead before we start; it's not a design option, it's a 100x budget violation. Same-region Redis is 0.5–1ms — affordable, but it puts a network hop and a dependency into every request. Reading a counter in the gateway's own memory is under a microsecond — effectively free. Those three numbers are the architecture: they force a tiered design where the common case stays local.
Memory. Token-bucket state is two numbers per key (token count, last-refill timestamp). Say 50M active keys × 3 endpoint classes × ~80 bytes ≈ 12GB. A small Redis cluster holds this comfortably. Memory is not the problem here. Coordination is — I'll say that out loud so the interviewer knows I've found the right fire.
The limiter exposes one call to the gateway code:
check(key, endpoint_class, cost=1)
-> { allowed: bool, remaining: int, retry_after_ms: int }
cost matters: one request can debit more than one token (a batch call costs 10; an LLM call debits tokens-of-text, more on that later). Policies live in a small config table, hot-reloaded to every gateway:
| Scope | Endpoint class | Rate | Burst | Mode | If limiter fails |
|---|---|---|---|---|---|
| per key | reads | 100/s | 200 | local | fail open |
| per key | writes | 10/s | 20 | strict | fail open |
| per key | auth (login, OTP) | 5/min | 5 | strict | fail closed |
| per tenant | all | 5,000/s | 8,000 | local | fail open |
The mode and failure columns are the interesting part of this data model — each is a committed decision I'll defend in the deep dives.
Quick recap from Part 1, then a commitment. Fixed window (count per calendar minute) has the boundary bug: 100 requests at 0:59 and 100 more at 1:01 — 200 in two seconds, all "legal." Sliding log (store every request timestamp) is exact but costs memory per request, which at 1M QPS is a non-starter. Leaky bucket smooths output to a fixed drip — great for protecting a fragile downstream, hostile to bursty clients. Token bucket: the bucket holds up to burst tokens, refills at rate tokens/sec, each request takes one. Steady rate enforced, short bursts allowed, and state is two numbers per key.
I'll commit to token bucket: it's burst-friendly, dirt cheap, and it's the industry default — Stripe and AWS both run their API limits on it. Cost of the choice: a client can legally spike to burst above steady rate, so downstream services must be sized for bursts. That's a price I want to pay, because the alternative is punishing every mobile client that wakes up and syncs.
A token bucket is a prepaid chai card that refills on a drip. The stall tops up your card at 100 rupees an hour, capped at 200. Sip steadily and you never notice the limit. Skip a while and you can splurge — but only up to the cap, and then you're back to the drip. Fixed-window counting is the stall that resets everyone's tab at the top of the hour: order 100 chais at 9:59 and 100 more at 10:01 and the tab never objects. The analogy breaks in one place: a chai card is one card in one pocket. Our problem is that the same "card" is being swiped at 40 stalls at once.
The architecture follows straight from the latency numbers: keep the common case in gateway memory, reconcile nearby, reconcile-globally lazily.
The check path (every request): gateway looks up the key's local bucket → refill by elapsed time → take a token or reject with 429. No network. The sync path (every 50ms): each gateway batches "key X spent N tokens" deltas for all its active keys into one pipelined call per Redis shard; the Lua script applies the debits to the regional bucket and returns the true remaining balance, which the gateway adopts as its new local ceiling. A key that overspent regionally comes back negative, and every gateway in the region starts rejecting it within 50ms.
Notice what the sync path does to load: instead of 1M Redis ops/sec (one per request), Redis sees ~20 batched syncs per second per gateway — hundreds of pipelined calls a second carrying thousands of deltas. The request path and the coordination path are fully decoupled. That decoupling is also where the accuracy leak lives, which brings us to the crux.
Exact global counting requires synchronous coordination on every request; the latency budget forbids synchronous coordination. Pick two of three: accurate, fast, distributed — you can't have all of them. Everything else in this chapter is bookkeeping around this triangle. The senior signal is not solving it — nobody solves it — it's quantifying exactly how much accuracy you're giving up, for which endpoints, and proving the leak is bounded and priced.
Walk the failure concretely first. Limit: 60/min. Two gateways, no sync. An abuser splits traffic evenly: each gateway counts 50, each says "50 < 60, allow." 100 requests hit the origin against a limit of 60. Local-only counters don't fail loudly — they under-count and over-admit, silently, by a factor of up-to-the-gateway-count. With 40 gateways, a 100/s limit quietly becomes a 4,000/s limit exactly when someone malicious — the very person limits exist for — spreads their load.
Two obvious fixes, both wrong for the general case, and I'll reject them out loud. Central counter: every request checks regional Redis synchronously — accurate within a region, but adds ~1ms and a hard dependency to every request, and at 1M QPS it means a Redis fleet doing 1M atomic ops/sec that, if it stumbles, takes the entire API with it. I'll use it selectively, not universally. Sticky routing: consistent-hash each key to one gateway, so its counter is exact and local. Clean on a whiteboard, breaks in production: geo/anycast routing means a global customer's requests genuinely originate on three continents; hashing them to one gateway means cross-region hops (there goes the latency budget) and turns every hot key into a hot machine. It's a fine design for a single-region internal API — worth saying, because knowing where a rejected design works is a senior signal.
So: local buckets with 50ms async sync, and now the honest math. Worst case over-admission ≈ per-key burst × number of gateways, for one sync window. Key limit 100/s, burst 200, 40 gateways, 50ms sync. An attacker who sprays all 40 gateways in the same 50ms can seed 40 fresh local buckets, each holding the full burst of 200: up to 8,000 requests admitted — 40x the burst — before the first sync lands. Within ~50ms, the regional bucket goes deeply negative, every gateway adopts the negative balance, and the key is choked until it repays at 100 tokens/s. So the leak is a one-time spike of burst × gateways, lasting at most one sync interval, followed by automatic payback.
Is 8,000 extra requests acceptable? Depends entirely on what they cost — which is why "mode" is a per-endpoint policy, not a global setting:
And there's a knob between the extremes: cap how much burst a local bucket may hold before checking in — say burst/8 per gateway — and the worst case drops 8x at the cost of hot keys syncing more often. I'd ship with that knob conservative and tune it from measurements.
A stadium with 40 entrance gates and a 60,000-person fire-safety limit. Give every gate its own clicker and radio a total every minute: fast lines, but between radio calls the gates are collectively blind — 40 gates × a burst of entries each can overshoot the cap until the next call, after which gates start turning people away. One central turnstile would count perfectly and create a two-hour queue. And the VIP vault door doesn't use clickers at all — the guard phones the control room and waits for a yes, because being wrong at that door is expensive. Clickers for the crowd, the phone for the vault: that's local mode and strict mode.
Cloudflare built rate limiting that enforces limits for millions of customer domains across its edge network, and they wrote up the design in "How we built rate limiting capable of scaling to millions of domains." Two decisions stand out. First, they chose not to count globally: counters live inside each PoP (point of presence), because anycast routing naturally pins a given client to its nearest PoP — so per-PoP counting is accurate for the case that matters, a real client or botnet node hammering from one place. Second, their storage was memcached, which offers only GET, SET, and atomic INCR — no way to run multi-step bucket logic atomically. So instead of a token bucket they used the sliding-window approximation: weight the previous window's count by how much of it still overlaps, add the current window. Two counters per key, no timestamps stored, and their measured error was about 0.003% across a 400-million-request sample. It's a beautiful example of the whole game: pick where counting is allowed to be local, then pick the algorithm your atomic primitives can actually support.
The regional tier has its own race to close. A bucket check is read-modify-write: read tokens, decide, write back. Two gateways syncing the same key concurrently can interleave: both read "1 token left," both decide yes, both write 0 — two spent, one existed. At thousands of syncs a second this isn't a corner case; it's a steady drip of exactly the over-admission we built this tier to stop.
Redis's fix is elegant: it's single-threaded for command execution, and a Lua script runs as one uninterruptible unit. Put the entire refill-decide-debit inside a script and the race cannot happen — no locks, no retries. The bucket, sketched:
-- KEYS[1] = rl:{acct_9f3k}:reads
local rate, burst = tonumber(ARGV[1]), tonumber(ARGV[2])
local now, cost = tonumber(ARGV[3]), tonumber(ARGV[4])
local s = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(s[1]) or burst
local ts = tonumber(s[2]) or now
-- refill for elapsed time, capped at burst
tokens = math.min(burst, tokens + (now - ts) / 1000 * rate)
local ok = tokens >= cost
if ok then tokens = tokens - cost end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], 120000) -- idle keys self-clean
return { ok, tokens }
Three operational details that interviewers probe. Sharding: 12GB of state and the sync load spread over a Redis Cluster, keys distributed by hash. The {acct_9f3k} braces are a hash tag — Redis Cluster hashes only the tagged part, so all of one account's keys (reads bucket, writes bucket, tenant bucket) land on the same shard and one script can touch them together. The TTL means idle keys evaporate; memory tracks active keys, not all keys ever seen. The hot key: one abuser hammering one key means all traffic for that key converges on a single shard — sharding doesn't help, because it's one key. Here the tiered design quietly saves us: the abuser's requests hit gateway-local buckets, which start rejecting locally once the synced balance goes negative. The shard never sees per-request traffic — only ~20 delta syncs/sec per gateway, bounded regardless of how hard the key is hammered. The local tier isn't just a latency optimization; it's a shield that turns a million-QPS attack on one key into a few hundred Redis ops per second. Cheap 429s are the point: a rejection served from gateway memory costs microseconds and touches nothing downstream.
Fail open or fail closed? When regional Redis is unreachable, the gateway must choose alone. I commit per endpoint class, by asking which mistake costs more. Reads fail open (local-only counting continues, accuracy degrades): blocking every legitimate read because the limiter — a protection layer — is sick would be the outage causing itself; over-admitting cached reads for a few minutes is absorbable. Auth endpoints fail closed (only the tight local allowance survives; anything beyond it is rejected): the 5/min limit on login and OTP attempts is a security control against credential-stuffing and brute force. Failing open there means every limiter outage is an attack window — and attackers can create limiter outages. A real user retrying a password twice stays under the local allowance and never notices; a bot spraying thousands of attempts hits the wall. Same mechanism, opposite defaults, each derived from the cost of being wrong.
Reply craft. A 429 is a conversation, not a door slam:
HTTP/1.1 429 Too Many Requests
Retry-After: 2
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1755162032
Well-behaved clients read Retry-After and back off; your SDKs should do exponential backoff with full jitter built in, because a thousand clients all told "retry in 2s" will return as one synchronized stampede — so jitter the server's own Retry-After values too. And X-RateLimit-Remaining quietly deflects support tickets: clients can watch their own budget drain instead of filing "API is broken" when they hit the wall.
Multi-tenant fairness. One bucket per key isn't enough when keys belong to organizations. I run two chained checks: the per-user bucket and the tenant-wide bucket, both must pass. The tenant bucket protects the platform from a tenant whose thousand users are collectively huge; the per-user bucket protects a tenant from its own noisiest user eating the whole org allowance. Chained buckets in one Lua script (hash tags put them on the same shard) keep it one round trip.
LLM APIs change the unit. For inference endpoints, requests/min is the wrong currency — one request can cost 100 tokens or 100,000, a 1000x cost range that request-counting can't see, which is why Anthropic and OpenAI publish tokens-per-minute limits. Same token-bucket machinery, but cost becomes the token count: debit an estimate at admission (prompt tokens + max_tokens), reconcile against actual usage when the response completes, refund the difference. Cost-proportional limiting like this is its own deep topic — p3c23 takes it on properly.
Stripe's engineering blog post "Scaling your API with rate limiters" (Paul Tarjan) remains the most-cited production writeup of this problem, and its lesson is that one limiter is never enough. Stripe runs four in concert: a token-bucket request limiter per user (the classic), a concurrent-request limiter capping in-flight requests, and two load shedders — one reserving fleet capacity so that critical calls like charge creation survive even when analytics-style traffic floods in, and a worker-utilization shedder as the last line, dropping low-priority traffic only as workers approach saturation. The request limiter is a token bucket, and the whole set runs on Redis. Two operational details are pure senior craft: they fail open — if Redis is having trouble, requests go through rather than the limiter becoming the outage — and they rolled the system out in shadow mode first, logging what would have been rejected for real traffic before turning on enforcement, then enabling it gradually behind feature flags. Rate limiting at Stripe isn't one guard at one door; it's layered defenses, each catching what the previous layer can't see.
Regional Redis cluster down. Gateways detect failed syncs and enter local-only mode: reads keep flowing with degraded accuracy (over-admission bounded by the math above), strict-mode endpoints follow their fail-open/fail-closed policy, auth stays clamped. I'll say the quiet part loudly: in this mode the limiter is lying a little by design, and the dashboard should say so — a visible "degraded accuracy" state, not silent drift. Recovery is clean because Redis state has TTLs: buckets rebuild from live traffic within a window or two; there is no replay or backfill to run.
What I monitor — each metric watches one specific way this design fails:
Rollout. Never enforce a new limiter cold. Shadow mode first — evaluate every request, log the would-be 429s, enforce nothing — because the first thing shadow data shows you is which legitimate heavy users your chosen limits would have broken. Fix the limits, then enforce endpoint class by endpoint class behind flags, cheap reads first. Limit changes go through config, take effect in seconds, and roll back the same way.
At 10x (10M QPS, ~400 gateways): the local tier scales for free — it's per-gateway memory. Two things break. The over-admission bound is proportional to gateway count, so it grew 10x: answer is shrinking the local burst slice and promoting more endpoint classes to strict, trading a little latency back for accuracy. And Redis sync fan-in grows with gateway count: answer is an aggregation hop — gateways sync to a per-AZ aggregator, aggregators sync to regional Redis — the same tiering trick applied once more.
This problem is a latency-vs-accuracy interrogation. A passing senior answer commits to tiered counting, quantifies the over-admission window unprompted, and splits policy by endpoint cost. Expect these probes:
| Decision | Why | Cost accepted |
|---|---|---|
| Token bucket | Burst-friendly, two numbers per key, industry default | Legal short spikes to burst above steady rate; downstream sized for it |
| Local buckets + 50ms async sync | Check stays in-process (~µs); limiter can never take the API down | Over-admission up to burst × gateways for one sync window |
| Strict mode for expensive endpoints | Mistake cost ≫ 1ms; these paths are slow and rare anyway | +1ms and a regional Redis dependency on those paths |
| Regional truth, lazy global reconcile | Cross-continent RTT (80–150ms) can't live in the request path | A key can briefly draw ~limit per region before squeeze |
| Fail open (reads) / fail closed (auth) | Priced by cost-of-mistake asymmetry per class | Outage → over-admission on reads, false 429s possible on auth |
| Redis Cluster + Lua + hash tags | Atomic read-modify-write; one shard owns each key's state | Abused key = hot shard, absorbed by local pre-filter |
A junior asks: "Why not keep one global counter per API key in a database? It'd always be correct." Explain the design in 5–6 sentences.
One global counter means every API request, worldwide, waits on a round trip to that counter — and between continents that's 100ms, on an API where our whole latency budget for limiting is one millisecond. So we keep a small token bucket in each gateway's own memory and answer instantly, then every 50ms all gateways report their spending to a shared Redis in their region, which corrects everyone's balance. The price is honesty about a small window: for up to 50ms, gateways can collectively let through more than the limit — at worst the burst size times the number of gateways — and then the shared balance goes negative and everyone starts rejecting until it's paid back. For cheap requests that leak is harmless, so we take the speed. For expensive things like payments we skip the shortcut and check Redis synchronously, because one extra millisecond is nothing on a path where being wrong costs real money. The trick isn't avoiding the accuracy-vs-speed trade — it's choosing it per endpoint instead of once for everything.
Add rate limiting for unauthenticated traffic (no API key — only an IP address). What changes, and what new problem does IP-keying create?
Cardinality explodes (billions of possible IPs vs 50M keys), so count only at the edge tier with aggressive TTLs — never let IP keys flood regional Redis. The nasty new problem is shared IPs: mobile carrier-grade NAT can put tens of thousands of innocent users behind one address, so a tight per-IP limit punishes a whole city for one bot. Answer: coarse per-IP limits (protect infrastructure, not fairness), local-only counting — this is exactly the case where Cloudflare's per-PoP argument holds, since a given IP's traffic arrives via anycast at one PoP — and push the real abuse decision to better signals (fingerprints, behavior scoring) upstream of the limiter.
The fleet grows from 40 to 400 gateways. Redo the worst-case over-admission math for a key with limit 100/s, burst 200, sync 50ms — and name the knob you'd turn first.
Worst case ≈ burst × gateways = 200 × 400 = 80,000 requests in one 50ms window — 400x the burst, clearly unacceptable even for cheap reads. First knob: cap the local burst slice, e.g. each gateway may hold at most burst/32 ≈ 6 tokens before checking in, cutting the bound to ~2,500. Second: shorten the sync interval for keys observed to be hot (adaptive sync). Third: promote more endpoint classes to strict mode. Also note the fan-in fix: 400 gateways syncing directly is Redis pressure, so insert per-AZ aggregators — tiering, applied again.
Requirement flip: sales now sells plans with a hard cap of exactly 1,000,000 calls per month, and overage is billed per call. Does your limiter handle this?
No — and saying so crisply is the senior move. The token-bucket tier is a protection system: approximate, fast, memory-resident, state that can evaporate on restart. A billed monthly quota is an accounting system: it must be durable, exact, auditable, and idempotent under retries (double-counting a retried request is a billing bug). Build it separately: an event stream of admitted requests into a durable, idempotent counter/ledger (p2c02), checked asynchronously — a key over its monthly cap gets flagged into the limiter's policy table within seconds. Latency-critical approximate limiting in the request path; slow exact accounting beside it. Two systems, honest seam.
A teammate proposes: "Consistent-hash every API key to one specific gateway. Counters become exact and local; delete the Redis tier." Write the two-sided critique.
Real strengths: exact counting, zero sync machinery, no Redis to operate. Real breakages at our scale: (1) a global customer's traffic enters on three continents — forwarding it all to the one "owner" gateway adds cross-region hops that blow the 1ms budget; (2) a hot key becomes a hot machine — one abuser can saturate a whole gateway, and the limiter's job was to prevent exactly that; (3) every deploy or autoscale event reshuffles the hash ring and resets counts mid-window. Verdict: sound for a single-region internal API with even key distribution; wrong for a global public API. Knowing which context flips the verdict is the point of the exercise.
Design the debit flow for an LLM endpoint limited at 100K tokens/min, where you must admit or reject before knowing how many tokens the response will use.
Reserve-then-reconcile. At admission, debit a pessimistic estimate: counted prompt tokens + the request's max_tokens (the contractual ceiling on output). Run inference. On completion, refund the unused difference to the bucket. Two edge cases prove you've thought it through: streaming responses that are cancelled mid-generation should reconcile at cancellation (charge actual tokens generated), and clients who set max_tokens huge "just in case" will self-throttle — their own reservations exhaust their bucket, which is a feature, not a bug: it prices honesty into the API. Use strict mode: GPU-seconds are exactly the expensive-endpoint class that over-admission must not touch. Full treatment in p3c23.
"Design a job scheduler — think cron, but distributed. Millions of jobs, a fleet of machines." The interviewer says it casually, like they're asking for a to-do list app. Don't be fooled. Hiding inside this cozy little prompt is one of the sharpest correctness questions in distributed systems, and the interviewer knows exactly where it's buried: what happens when the machine that fires the job dies halfway through firing it?
Feel the stakes first. A job scheduler is the thing that runs "charge every subscriber on the 1st," "send the daily digest at 9am," "rotate the TLS certificates before they expire." If it fires a job zero times, someone's certificate expires and half the internet gets a scary browser warning. If it fires the same job twice, a customer gets charged twice, and now you're trending on social media for the wrong reason. Zero and two are both catastrophic, and the entire problem is that a distributed system makes "exactly one" the single hardest number to hit.
Mid-level candidates draw a cron box with an arrow to some workers and start talking about queues. Senior candidates walk straight at the fire: "the crux here is making sure each due job is fired by exactly one scheduler, and being honest that once it's fired, exactly-once execution is impossible — so I'll design for at-least-once plus idempotency." Say that in the first ten minutes and the interview changes temperature. Let's earn the right to say it.
"Here's what I'll build: users register jobs — either recurring on a cron schedule (0 9 * * * = 9am daily) or one-shot at a specific time. A job's payload says what to run: a queue topic plus arguments that some worker knows how to execute. I'll support retries with backoff, priorities, and a per-job misfire policy — what to do when we're late, because one day we will be. I'll sketch DAG dependencies — a directed acyclic graph, meaning job B runs after job A succeeds and no arrow ever loops back, Airflow-style — as an extension, not the spine. Out of scope: the workers' business logic itself, and a UI. Fair?"
Non-functional, with numbers I'm committing to rather than fishing for:
Notice the asymmetry I just committed to: a missed fire is a silent lie that pages nobody, so the design must make misses structurally impossible to hide. A duplicate fire is loud and survivable if — only if — job authors are given the tools to make their jobs idempotent. That asymmetry drives everything below.
150M firings/day is about 1,700 per second — on average. And the average is a lie, because humans don't schedule uniformly. Look at any real crontab collection: schedules cluster brutally on round numbers. 0 * * * *, */5 * * * *, and the king of them all, 0 0 * * * — midnight. The load isn't a stream; it's a heartbeat. Second 0 of every minute carries a spike, minute 0 of every hour carries a bigger one, and midnight UTC is a tidal wave. With 10M jobs, it's realistic that hundreds of thousands of firings all come due in the same second at 00:00 UTC.
Say the conclusion out loud, because this number just designed half the system: "I cannot build this with a timer per job — 10M in-memory timers that vanish on crash — and I cannot build it to dispatch the midnight wave instantaneously. So: due jobs live in a database indexed by next-fire-time, a scanner tier sweeps that index every second in batches, and a queue absorbs the wave so workers drain it over a few seconds. My 5-second p99 SLO is exactly the size of that drain." Numbers → architecture. Also storage: 10M job rows at ~1 KB is 10 GB — a rounding error, one Postgres. But 150M run-history rows a day at ~500 bytes is 75 GB a day — so run history gets a partitioned table with a 30-day retention and archival to object storage, and it never shares a disk budget with the job specs. The tiny table is the brain; the huge table is just a diary.
Google hit every part of this problem and wrote it up as a chapter of the Site Reliability Engineering book: "Distributed Periodic Scheduling with Cron." Their internal cron service replaced the single trusty cron daemon with a small group of replicas running Paxos — the leader replica is the only one allowed to launch jobs, and before it launches anything it announces "I am about to start this launch" through consensus, so a new leader after a crash knows exactly which launches were in flight. Even Google couldn't dodge the two-generals trap: for a launch caught mid-flight during failover, the new leader cannot tell whether that job already started, so it has to pick a side — relaunch and risk a duplicate, or skip and risk a miss. Their published rule is one sentence long: they favor skipping launches rather than risking double launches, as far as the infrastructure allows. Exactly-once was never on the menu; a stated default was. And the thundering herd was real enough that they extended crontab syntax itself: a ? in a schedule field means "any value is fine — you pick," and the system hashes the job's own config across the allowed range to choose one, deterministically smearing the midnight mob over the whole hour. When a company with Paxos in its toolbox concludes "exactly-once is impossible, spread the herd, and write down which way you lean on skip-versus-rerun," believe them.
POST /jobs — {name, schedule: "0 9 * * *" | run_at, payload: {queue, args}, priority, max_attempts, timeout_s, misfire_policy, jitter_s?} → 201 {job_id}. Accepts an idempotency key so a retried create doesn't register two jobs (chapter 2.2 habit).PATCH /jobs/{id} — pause, resume, edit schedule. DELETE to remove.POST /jobs/{id}/run-now — manual trigger; also the hook the DAG layer will use later.GET /jobs/{id}/runs — execution history and current state.Two tables carry the system. jobs: id, schedule, payload, priority, max_attempts, timeout_s, misfire_policy, enabled — and the load-bearing column, next_run_at, materialized at write time. When you register 0 9 * * *, we parse the cron expression once and store the concrete next timestamp. An index on (next_run_at) WHERE enabled turns "what's due?" into a cheap range scan off the front of the index — no cron parsing on the hot path, ever. runs: run_id, job_id, scheduled_for, state (pending → running → succeeded | failed | dead), attempt, claimed_by, lease_expires_at. Hold on to runs — it's about to do three different jobs at once.
Write path: client registers a job → API parses the cron expression, computes next_run_at, inserts the row. Done. Registering a job is boring, and that's a feature.
Fire path (the real read path): every second, each scanner asks the job store for a batch of due jobs and claims them (deep dive 1 is entirely about this word). For each claimed job, in the same transaction: insert a runs row in state pending, and advance the job's next_run_at to its next cron occurrence. Commit. Then enqueue each run_id onto the right priority queue. Workers pull a run_id, flip the run to running with a lease, execute the payload, report success or failure. The runs row inserted inside the claim transaction is doing outbox duty (the chapter 2.1 pattern — record the intent to send in the same transaction as the state change, so a crash between them is impossible): if a scanner commits its claim and then crashes before enqueueing, a sweeper finds pending runs older than a minute and re-enqueues them — the firing can be delayed by a crash, but never silently lost. And if the sweeper re-enqueues something the scanner actually did enqueue, no harm: the worker's claim on the run is conditional (UPDATE runs SET state='running' … WHERE run_id=X AND state='pending'), so the second delivery finds a non-pending run and drops dead. Queue-level duplicates: neutralized at the run level, for free.
For years, a meaningful slice of Slack ran on what their engineers frankly called the "cron box": one node holding a copy of every cron script and one giant crontab with all the schedules. As Slack grew, the scripts multiplied into the thousands and each one processed more data, so the fix was always the same — buy a bigger box. More CPU, more RAM, repeat. It worked until it didn't: a single misprovisioned or misconfigured node could take down chunks of user-visible Slack functionality, because a single machine was the company-wide scheduler. The replacement they described in 2023 has the same skeleton as the whiteboard design in this chapter: a Go service called the Scheduled Job Conductor, running on Bedrock (their Kubernetes wrapper), decides what's due; execution is handed off to their existing job queue platform instead of running locally; and a Vitess-backed table does deduplication and run tracking — the same "database as arbiter of what actually fired" idea we're about to deep-dive. They diverge from this chapter on one point, and it's worth noticing: their conductor pods elect a leader, one scheduling while the rest stand by, where we'll let the database referee instead. The lesson isn't that cron is bad; it's that "one box fires everything" has no next step except a bigger box, and the escape route is separating deciding to fire from executing the work.
Clock check: the loop is closed — register, fire, execute, report, recover. Now I'd tell the interviewer my agenda: "Three things deserve depth: who gets to fire a due job, what happens when a worker dies mid-job, and the time policies — misfires, priorities, DAGs. First one first, it's the crux."
This is what the problem was invented to probe. You need multiple scanners — one is a single point of failure, and the midnight wave needs parallel hands. But two scanners running SELECT … WHERE next_run_at <= now() in the same second see the same due jobs, and both will happily fire them. Every duplicate-charge horror story starts here. You must make "claim this due job" an atomic, arbitrated act — and you must know what happens when a claimer dies holding claims. Interviewers who ask this problem are mostly grading this one answer: naive scan = fail, global lock = mediocre, row-level claim or leader+fencing, argued honestly = senior.
The committed answer: let the database arbitrate, with SELECT … FOR UPDATE SKIP LOCKED. Every scanner runs the same loop, once per second:
BEGIN;
SELECT id, payload, schedule FROM jobs
WHERE enabled AND next_run_at <= now()
ORDER BY next_run_at
LIMIT 500
FOR UPDATE SKIP LOCKED;
-- for each row: INSERT runs(run_id, job_id, scheduled_for, 'pending');
-- UPDATE jobs SET next_run_at = next_occurrence(schedule);
COMMIT; -- then enqueue the run_ids
Unpack the two magic words. FOR UPDATE means: lock the rows I selected until my transaction ends — nobody else can touch them. SKIP LOCKED means: if a row is already locked by someone else, don't wait for it and don't error — pretend it isn't there and take the next one. Put together, ten scanners hitting the same due-list don't fight: the first grabs rows 1–500, the second silently skips those and grabs 501–1000, and so on. The scanners never talk to each other, hold no state, and don't even know how many siblings they have — the database's lock manager is doing the coordination, and it's had twenty years of production hardening at exactly this. Deploying more scanners is just more hands pulling from the same rail.
It's the ticket rail in a restaurant kitchen. Orders (due jobs) hang on the rail; any free cook grabs the next ticket. The grab is the whole protocol — a ticket in one cook's hand physically cannot be in another's, no head chef assigns work, and adding a cook at rush hour needs zero reorganization. And when a cook faints holding three tickets? In our kitchen the tickets snap back onto the rail on their own: an uncommitted transaction's locks vanish when its connection dies, so the crashed scanner's claimed rows roll back to due-and-unlocked, and the next sweep picks them up. The analogy breaks in one place, and it's the honest part: a fainted cook's ticket comes back even if the dish was already cooked — which is why the next deep dive exists.
Walk the crash cases, because the interviewer will: scanner dies mid-transaction → rollback, rows unlock, nothing fired, next tick refires — a few seconds of delay, zero loss, zero duplicates. Scanner dies after commit, before enqueue → next_run_at already advanced so no refire, and the pending run row is the outbox that the sweeper turns into a (possibly duplicate, provably harmless) enqueue. There is no crash point that silently loses a firing — that's the property the requirements demanded, and I can point to the line of SQL that provides it in every case.
The alternative I'm consciously rejecting: a leader-elected scheduler tier — an etcd lease (chapter 2.3) picks one scanner as the only firer; followers idle. Google's cron works this way with Paxos. It's the right call when the "database" is itself the thing you're building, or when firing decisions need complex in-memory state. But it buys me real problems: a failover gap where nothing fires, a single leader as throughput ceiling for the midnight wave, and the classic stale-leader trap — a leader stalls in a garbage-collection pause, its lease expires, a new leader takes over, then the old one wakes up and fires everything it was about to fire. Chapter 2.3 taught the fix — fencing tokens checked at the point of write — but notice that the fence has to live in the job store anyway. If the store must arbitrate regardless, let it arbitrate with row locks and skip the election entirely. At 1,700 claims/sec average, a single Postgres does this without noticing; I'll name the ceiling honestly in the ops section.
For its first six years, Apache Airflow — the most widely deployed job orchestrator in data engineering — had exactly one scheduler process. It was the project's most notorious weakness: if the scheduler crashed, every DAG in the company stopped firing until someone restarted it, and its scan loop was a throughput ceiling teams hit constantly. When Airflow 2.0 shipped in December 2020, the headline feature was the HA scheduler, and the mechanism is precisely this deep dive's committed answer: run as many scheduler replicas as you like, all active, no leader election at all — each one claims work through database row-level locks using SELECT … FOR UPDATE SKIP LOCKED, so the metadata database arbitrates and locked rows are invisibly skipped by peer schedulers. The dependency is real enough that databases lacking SKIP LOCKED made the feature unusable — MariaDB didn't implement it until 10.6, and Airflow users running multiple schedulers on older MariaDB hit deadlocks and duplicate scheduling. A decade of orchestrator experience distilled into a design: don't elect a chief, let the database referee.
Now the honesty section. A worker pulls run #881: "charge customer 4412 their $9.99." It charges the card. Then, before it can report success, the machine dies. The run row still says running, the lease expires, and the system now faces a question it cannot answer: did the charge happen? This is the Two Generals problem from chapter 2.2, wearing work clothes: when a machine goes silent, the rest of the system cannot distinguish "did the work, died before telling us" from "died before doing the work." No protocol fixes this. No amount of acks fixes this — the last ack can always be the one that's lost.
So you choose which failure you'd rather have. Re-run on silence and you get at-least-once — duplicates possible. Never re-run and you get at-most-once — losses possible. For a scheduler, I commit to at-least-once: a missed payroll run is a silent disaster, a duplicated attempt is survivable if jobs are idempotent. And then — this is the part that separates senior answers — I make idempotency a first-class, documented obligation of every job author, not a footnote. The scheduler's contract reads: "your job will occasionally be started twice for the same scheduled firing. Build accordingly."
The mechanics that keep duplicates rare (never zero — rare):
lease_expires_at = now() + 90s. Long jobs heartbeat every 30 seconds, each beat extending the lease — a 6-hour job holds its claim through 720 renewals. A worker that stops heartbeating is presumed dead; that's the same visibility-timeout idea SQS uses, done in our own table.state='running' AND lease_expires_at < now(), flips them back to pending with attempt+1, and re-enqueues. Retries use exponential backoff with jitter (chapter 2.7's retry hygiene: a failing dependency doesn't need 10,000 synchronized retries).max_attempts, the run goes to state dead — our dead-letter queue — and alerts a human. A DLQ nobody watches is a landfill; the DLQ arrival rate is a first-class alarm, not a dashboard curiosity.The idempotency cookbook — what "build accordingly" concretely means, straight from chapter 2.2's toolbox, published as the scheduler team's most-read doc:
(job_id, scheduled_for) — not by attempt, not by run_id. All retries and duplicates of one scheduled firing share this key; two legitimate firings (today's 9am vs tomorrow's) don't.INSERT … ON CONFLICT DO NOTHING keyed on the firing; UPDATE invoices SET state='paid' WHERE id=? AND state='due' — the second execution matches zero rows and dies quietly.report.tmp.{run_id}, atomically rename to report_2026-08-14. Duplicate runs overwrite with identical bytes instead of appending garbage.The lease-and-heartbeat dance is a library loan with renewals. You check out a book (claim the run), and as long as you keep renewing (heartbeating), the library knows you're still reading and leaves you alone — for months, if needed. Stop renewing and the library doesn't know if you lost the book or just lost interest; after the grace period it declares the copy gone and orders a replacement (re-enqueues the run). And here's the honest edge of the analogy: if you then stroll in and return the original, the library owns two copies. That's the duplicate. The library can't prevent it — it could only have waited longer, which just delays every replacement for every genuinely lost book. Idempotency is the shelf accepting the second copy without creating a second catalog entry.
Misfire policy. The scheduler tier was down for 10 minutes — deploy gone wrong, database failover, whatever. It's 09:10 and the index holds thousands of jobs stamped 09:00. What now? There is no universal answer, which is exactly why it's per-job config with three values. RUN_NOW: fire late, once — right for a daily digest; 9:10 is fine. SKIP: recompute next_run_at, catch the next occurrence — right for high-frequency work like a */5 cache refresh, where a late run is worthless because a fresh one is 3 minutes away. CATCH_UP: fire once per missed occurrence — right only for jobs whose payload is parameterized by the period, like "generate the statement for hour X"; skipping leaves a hole in the ledger. My committed defaults: RUN_NOW if we're within one schedule interval of the miss, SKIP beyond that, CATCH_UP strictly opt-in — because an accidental catch-up after a long outage is a self-inflicted thundering herd of stale work. (Quartz calls this the misfire threshold; every mature scheduler grows this knob, so put it in the API on day one.)
Priorities. A certificate-rotation job must not queue behind 200,000 bulk thumbnail re-encodes. Committed design: three queues — critical, default, bulk — with dedicated worker pools, not one pool doing strict priority-order polling. Strict priority starves: a busy critical queue means bulk work waits forever, and "forever" always eventually includes something that mattered. Dedicated pools guarantee every tier a floor of capacity; critical's pool is simply provisioned so its queue depth stays near zero.
DAG dependencies — one section, deliberately not the spine. "Run build_ledger after pull_orders, pull_refunds, and pull_fx_rates all succeed" — the Airflow shape. Layer it above the core: a dependencies table, and a completion event for every finished run. When a parent succeeds, atomically decrement the child's unmet-dependency counter; the decrement that hits zero triggers the child through the same run-now path a human would use. Fan-in is just "counter reaches zero exactly once" — the decrement must be a conditional DB update, not a read-then-write, or two parents finishing simultaneously both see 1 and neither fires the child. Validate at submission time with a depth-first search and reject cycles outright — A→B→A is a deadlock you detect at write time for free or debug in production at great expense. The senior note: the core scheduler stays completely unaware of DAGs; it fires jobs, and one of its consumers happens to be a dependency engine. That layering is why the core stays testable.
build_ledger — exactly one trigger even when parents finish in the same millisecond.What I monitor — the four numbers that describe this system's health: (1) fire-time skew p99: scheduled_for versus worker pickup — my 5-second SLO, measured, graphed per priority tier. (2) Queue depth per tier — critical's depth should hug zero; a rising default-queue depth is the earliest backpressure signal (chapter 2.7). (3) DLQ arrival rate — a step change means some job's dependency broke. (4) Oldest unclaimed due job — now() - min(next_run_at) over enabled jobs. This one deserves a sentence: if the whole scanner tier silently wedges, there are no errors, no failed runs, no bad metrics — nothing happens, which is the one failure a normal error-rate alarm can't see. The oldest-due-job gauge climbing is how silence becomes visible. Belt-and-suspenders: a dead-man's-switch canary job fires every minute and pings an external monitor; the monitor alarms on absence. The scheduler must never be the only witness to its own death.
Clock skew, honestly. Here's an underrated virtue of the SKIP LOCKED design: correctness doesn't depend on scanner clocks at all. The claim query's now() evaluates on the database's clock — one authoritative clock, however wrong, applied consistently. Scanner clock skew affects nothing; NTP keeps worker clocks within tens of milliseconds for lease arithmetic, and our tolerances are in seconds. Contrast Quartz, the JVM world's veteran scheduler: its clustering mode coordinates through database row locks too, but each node decides due-ness on its own clock, and its documentation is blunt about the consequence — never cluster across machines unless a time-sync daemon keeps every node's clock within a second of the others. Node clocks are load-bearing there; in our design they aren't, and that's a dependency worth not having. When a twenty-year-old scheduler puts a clock warning in its config docs, put the sentence in your design: "one authoritative clock for due-ness; everything else tolerates seconds of skew."
Rollouts. Scanners and workers are stateless — rolling deploys, no drama. The scary change is the cron parser: a subtle parsing bug rewrites next_run_at wrong for millions of jobs on their next fire. So parser changes ship in shadow mode first: new code computes next occurrences alongside old, we diff a day of results, deploy only on zero diffs. Cheap, and it converts the worst silent-corruption risk into a boring report.
What breaks at 10x? 100M jobs, 17K firings/sec, and the answer is: the claim path, first and specifically. One Postgres doing tens of thousands of row locks per second with every scanner contending on the same index range stops being comfortable. The fix is the one we learned in chapter 1.9: partition. Hash job_id into 16 logical shards (a shard column, or 16 physical DBs later); each scanner sweeps its assigned shards; claims never contend across shards, and the midnight wave splits 16 ways. Second-order casualty: the runs firehose — 100M+ inserts/day — which is why it was born partitioned-by-day with an archival policy. Name the ceiling before the interviewer does: "single-store claiming carries us to a few thousand claims/sec; past that I shard the claim, and the design doesn't change shape — it multiplies."
This problem is a correctness interview wearing an infrastructure costume. The interviewer is listening for one sentence above all: "exactly-once firing needs an arbiter; exactly-once execution is impossible, so at-least-once plus idempotency." Expect these probes, near-verbatim:
| Decision | Chosen | Cost accepted | Rejected alternative — and when it wins |
|---|---|---|---|
| Fire arbitration | DB row claims, FOR UPDATE SKIP LOCKED | All coordination funnels through one store; ~few-K claims/sec ceiling before sharding | Leader-elected tier (etcd lease + fencing) — wins when firing needs rich in-memory state, or there's no shared SQL store |
| Execution guarantee | At-least-once + mandatory idempotency contract | Duplicate runs happen; every job author carries an idempotency obligation | At-most-once — wins for cheap, loss-tolerant, non-idempotent-able work (some notifications) |
| Due-job detection | 1-second batch scan on next_run_at index | Up to ~1s inherent firing latency + queue drain time | Per-job in-memory timers — wins only for a small hot set needing sub-second precision (see exercise 1) |
| Priorities | Per-tier queues + dedicated worker pools | Some idle capacity reserved in quiet tiers | Single queue, strict priority — wins at tiny scale; starves bulk work at ours |
| Misfire default | RUN_NOW within one interval, else SKIP; CATCH_UP opt-in | Occasional late runs; missed occurrences for frequent jobs | Universal CATCH_UP — wins never as a default; it turns every outage into a replay storm |
| Job store | Single Postgres, runs partitioned by day | Vertical ceiling; deliberate re-shard milestone at ~10x | Sharded from day one — wins if you're Google-sized on day one, which the numbers say we aren't |
A junior teammate asks: "Why can't our scheduler just guarantee my job runs exactly once? Kafka has exactly-once, right?" Answer in five or six plain sentences.
Imagine your job charges a credit card, and the worker dies right after the charge but right before it tells anyone. From the outside, "did the work, died before reporting" and "died before doing anything" look identical — silence. Whatever the scheduler does next is a guess: re-run and it might duplicate the charge; don't re-run and it might have never happened. That's the Two Generals problem, and no protocol escapes it — even Kafka's "exactly-once" only covers effects inside Kafka's own transactional world, not your call to a bank's API. So we promise at-least-once — we will never silently lose your firing — and in exchange you make your job safe to repeat: key your side effects by (job, scheduled time) and write conditionally, so a second run finds the work done and does nothing. Duplicate attempts are our problem; duplicate effects are yours, and the cookbook makes them cheap to prevent.
Now add a requirement: a premium tier needs firing precision of ±100 ms (e.g., market-open triggers at 09:30:00.000). Your 1-second scan plus queue drain can't deliver that. What do you add, and what stays the same?
Don't rebuild the spine — add a small express lane. A pre-fetcher pulls the next 2 minutes of premium jobs into scheduler memory ahead of time (they're a tiny fraction of 10M), and a dedicated firer arms real in-memory timers and dispatches directly to a reserved worker pool, skipping the general queue. The claim still happens through the same DB transaction — arbitration and crash-recovery rules don't change, and it's still at-least-once (the pre-fetcher can crash and refire; premium jobs still need idempotency). You've traded generality for precision on a small hot set. This is also the honest answer to "why not timers for everything": 10M armed timers that evaporate on every crash and need re-arming from the DB anyway is the scan loop with extra steps — timers earn their complexity only when a small set needs precision the scan can't give.
10x it: 100M jobs, 17K firings/sec sustained, and a midnight wave of 3M due firings. Walk the failure order: what saturates first, second, third — and the fix for each.
First: the claim path — tens of thousands of contended row locks/sec on one index range in one Postgres. Fix: hash-partition jobs into 16+ shards, scanners sweep disjoint shards, claims stop contending; midnight splits 16 ways. Second: the runs table — 100M+ inserts/day plus status updates plus reaper scans. Fix: it's already partitioned by day; move status-history reads to replicas, batch status writes, archive aggressively. Third: worker capacity at midnight — 3M firings against a fixed fleet means the drain time, and thus fire-time skew, blows the SLO once a day. Fix: schedule-smearing (jitter for jobs that opt in, Google's ? trick as a nudge in the API), plus autoscaling the bulk pool ahead of known waves. Note what never changed: the claim transaction's shape, the lease protocol, the idempotency contract. Scale multiplied the design; it didn't reshape it — that's the sign the original commit was right.
Flip a requirement: a new job class — promotional push notifications — where a duplicate is worse than a miss (double-ping users and they uninstall). The PM asks for "at-most-once" jobs. What changes in the run lifecycle?
Invert the order of record-and-execute. Today: execute, then mark done — silence triggers re-run. For at-most-once: the worker marks the run completed (or "burned") before executing, and the reaper never re-enqueues this class — a lease expiry moves it to a unknown state for human review instead of retry. Now a crash between the mark and the send means a notification that never went out and never will: that's the miss you signed up for. Be the senior in the room about the boundary: this is at-most-once attempted delivery, not a duplicate-proof system end to end — if the push provider itself retries internally, users can still see doubles; true suppression needs an idempotency key at the provider. Misses are silent by design, so pair the class with delivery-rate monitoring, or you'll discover a month of missed promos from a revenue chart.
What breaks if… the whole scanner tier deadlocks on a bad deploy at 02:00 — processes alive, health checks green, zero jobs firing. Your error-rate alarms stay silent (nothing is running, so nothing fails). Design the detection.
Two independent tripwires. Inside: an "oldest unclaimed due job" gauge — now() - min(next_run_at) over enabled jobs — exported by a monitor that isn't the scanner; it climbs monotonically the moment firing stops, alert at 60s. Outside: a dead-man's switch — a canary job fires every minute and hits an external monitoring service, which alarms on absence of the ping; this catches the case where the database, the metrics pipeline, or the whole region went with the scanners. The principle worth saying in an interview: for a system whose job is to act, the deadly failure mode is silence, and silence never trips an error-rate alarm — you must alert on the absence of expected events, with at least one watcher outside the system's own blast radius.
A user submits DAG edges: ledger→report, report→cleanup, cleanup→ledger. When and how does your system reject this, and what exactly goes wrong if you accept it?
Reject at submission, inside the write transaction: load the affected component of the graph, run DFS from the new edge; finding a back edge (cleanup→ledger closes the loop) returns 400 cycle detected: ledger→report→cleanup→ledger — naming the path, because a bare "invalid" for a 200-node graph is user-hostile. Accept it and nothing crashes — that's the trap: each node's unmet-dependency counter (1) never reaches zero, so the cycle sits eternally pending. No errors, no retries, no DLQ entries — just three jobs that never run and a "why didn't cleanup fire?" ticket weeks later. It's the missed-fire problem in a new costume, invisible to error-based alerting; a "pending > 24h" watchdog is the backstop, but write-time DFS costing milliseconds is the actual fix.
next_run_at in the job store, index it, and batch-scan every second — the index range scan is the entire "what's due" algorithm.SELECT … FOR UPDATE SKIP LOCKED makes the database the fire arbiter: concurrent scanners claim disjoint rows, and a crashed claimer's locks vanish with its transaction — Airflow 2.0 ships exactly this instead of leader election.(job_id, scheduled_for), conditional writes."Design the trending list," the interviewer says. "Top ten hashtags on the home page. Last hour, last day. Refreshed while the user watches it."
It sounds like the easiest question in the gauntlet. Count things, sort, take the first ten — you've written that function in four lines on a laptop. So you say the obvious thing out loud: a hash map from hashtag to counter, sort the entries, slice the top ten.
Then the interviewer asks the only question that matters: "How big is that hash map?" And you start doing arithmetic on the whiteboard and discover that the map is a quarter of a terabyte, that it has to exist sixty times over because the window slides, that a single spam wave can double its size on purpose, and that the sort you casually mentioned is a sort of three hundred million entries happening every few seconds.
That's the whole problem. Top-K is a memory problem wearing a counting problem's clothes, and the way out is to stop insisting on exact answers in the hot path. You'll build a machine with two brains: a fast one that is wrong on purpose, in a direction you control, and a slow one that is exactly right and arrives hours late. Then you'll merge them — and be able to say precisely who gets hurt by the wrongness, because that sentence is what gets you the level.
"Trending" is the vaguest word in the interviewer's prompt and I want it nailed down in the first two minutes, because two very different products hide inside it.
Functional, committed:
Non-functional, with numbers I'll be held to:
Out of scope, said out loud so nobody thinks I forgot: personalized trends, ranking ML, deep spam classification. I will cover cheap anti-gaming, because a trending list with no bot defense is a billboard you hand to spammers.
That "exact for money" line isn't a detail. It's the seam the whole architecture splits along.
Three calculations, each of which changes the design. I do them out loud.
1. Volume. 10M events/sec peak, ~100 bytes each: 1 GB/sec at peak, and about 43 TB of raw events a day at the 5M average. Conclusion: nothing in the hot path may store an event. Whatever counts, counts in one pass and throws the event away. The raw stream survives only as a log in object storage for the batch path, where a terabyte is cheap.
2. Distinct keys — the killer. How many distinct video IDs, hashtags or search strings appear in one minute of 600M events? The head is small, the tail enormous: fresh uploads, one-off searches, typos, spam. Call it 50 million distinct keys a minute. Price the exact map: a key (16–30 bytes), a counter, plus pointers, load-factor slack and object headers. Eighty bytes an entry is optimistic. One minute = 50M × 80 B = 4 GB.
But I need a sliding hour — at 10:31 the answer covers 09:31 to 10:31 — so I keep per-minute maps and add them: 60 × 4 GB = 240 GB of hot RAM, for one scope. Add per-country scopes and it's a terabyte. Garbage-collected. Replicated, so double it. And the cruel part: a spammer minting 500 million junk hashtags an hour raises my memory bill on demand, by writing a for-loop.
Now the alternative. A count-min sketch (chapter p2c10) with 262,144 counters per row and 4 rows is 4 MB per minute; a 60-minute ring is 240 MB per shard, under 8 GB across 32 shards. 30x less memory — and, far more importantly, that number doesn't move when the distinct-key count moves. A billion junk keys cost the sketch zero extra bytes.
That constant is the decision. Not "sketches are cool" — "exact costs 240 GB and scales with an attacker's imagination; the sketch costs under 8 GB across the fleet and doesn't."
3. CPU and shard count. A sketch update is 4 hashes and 4 random memory writes — ~300 ns once cache misses are counted, so about 3M sketch updates/sec on one core. The sketch isn't the whole cost of an event, though: decoding the message and the heap check dominate it, so call the end-to-end work 1 microsecond, or 1M events/sec per core. At 10M/sec that's 10 cores pinned flat, and nobody runs a core at 100% — so about 20 busy cores. I'll commit to 32 shards (power of two, comfortable headroom). Each handles ~300K events/sec, about 18M per minute-bucket. Hold onto that 18M — it sets the error bound later.
You're at a stadium exit counting which city each departing fan came from, and you have 100 clipboards for 40,000 cities. So you assign cities to clipboards by a rule — first letter, say — and several cities share each clipboard. Any clipboard's tally is now "at least" the true count for any city on it, never less. Do it four times with four different rules, and for a given city take the smallest of its four tallies — the least contaminated view you have. You never undercount; you sometimes overcount someone sharing with a big city. That's a count-min sketch, and the bill is 100 clipboards however many cities walk past. Where it breaks: real clipboards let you read off which cities were counted. The sketch can't list its keys — you must bring it one to ask about, which is why it always travels with a heap.
One public read endpoint. The interesting part is the two fields most candidates leave out.
GET /v1/trending?scope=global&window=1h&metric=velocity&k=10
→ {
"as_of": "2026-08-14T10:31:20Z",
"window": "1h",
"estimated": true,
"items": [ {"rank":1, "key":"#monsoon", "count":812400}, ... ]
}
as_of and estimated exist because I promised staleness and approximation, and an API that hides its own guarantees becomes a support ticket six months later. Callers needing exact numbers use a different endpoint that reads the warehouse table and refuses to answer for the still-open hour.
Data shapes, all boring on purpose:
{key, key_type, user_id, country, ts}.lb:{scope}:{window}:{metric} → a small JSON document of the top 1000, rewritten every 5 seconds, TTL 60 seconds so a dead writer eventually shows as stale rather than lying forever.hourly_counts(hour_start, scope, key, count), partitioned by hour. This is the exact record, and it is what payouts read.key_baseline(scope, key, ewma_rate, updated_at).Write path. A view fires an event. The collector does one cheap thing before publishing: it keeps a small in-process map for a second and folds repeats together, so a viral video pulling 100,000 views a second through one collector becomes one message saying +100000. This pre-aggregation (Flink calls it local-global aggregation) is a 10–100x reduction on exactly the keys that would otherwise hurt, and it costs a hash map with a one-second lifetime.
Events go into Kafka partitioned by hash of the key, not round-robin — a real decision with a real cost, defended in the merge section. The short version: it makes each shard's answer about a key complete, which makes the merge trivially correct.
Each of 32 shard workers consumes one partition and holds a ring of per-minute count-min sketches plus a per-minute min-heap of its local top keys. Every minute it seals a bucket and ships a small summary up. The aggregator merges those into global and per-country lists per window and writes the blob to Redis every 5 seconds — event to published leaderboard in under 10 seconds.
Read path. The app reads one Redis key, or more often a CDN copy of it with a 5-second TTL. A million reads per second against a single blob is a textbook hot key (chapter p2c06), and the fix is the boring one: cache it close to users, serve stale while revalidating, never let a read reach the pipeline.
The slow path. The same Kafka stream is written to object storage as an immutable log. Hourly, a Flink or Spark job reads the closed hour and computes exact counts per key per scope into hourly_counts. Those rows serve analytics and payouts, and retroactively correct the leaderboard for closed hours.
Modern taste says kappa: one streaming pipeline, and replay the log through the same code to fix history. I like kappa and use it elsewhere — chapter p2c09's default. Here I'm choosing the lambda-flavored two-path design, and the justification is specific.
Replaying the fast path does not produce exactness, because the fast path is lossy by construction. The sketch already threw away what exactness needs; running the same approximate code over the same events again gives the same approximation. The batch path isn't a re-run — it's a different algorithm, one that can afford exactness because it may spend minutes and disk instead of microseconds and RAM. Second reason: late events. The stream closes a bucket after a couple of minutes of allowed lateness, and a phone that was offline for twenty minutes flushes afterwards — only the batch job sees the complete hour.
And the flip side, before the interviewer asks: if exactness weren't a requirement, I'd delete the batch tier. Sketches alone would serve the leaderboard fine. That second codebase earns its operational cost only because "exact for money" is real.
Twitter hit this exact wall and productized the answer. They open-sourced Summingbird in 2013 and published the design at VLDB 2014 (Boykin, Ritchie, O'Connell and Lin): you write an aggregation once, in a MapReduce-shaped Scala DSL, and the same code compiles to two runtimes — Scalding on Hadoop for the batch view, Storm for the real-time view — with reads merging the two views into one answer. The paper's motivation is the seam this chapter is built on: the batch layer is accurate but hours behind, the online layer is immediate but approximate and failure-prone, and the merge gives you both. The constraint that makes it work is algebraic: every aggregation must be a monoid — associative, with an identity — because partial results computed on different machines, in different time windows, must combine to the same answer regardless of order. Twitter's companion library Algebird ships monoid implementations of exactly the structures in this chapter: count-min sketch, HyperLogLog, Bloom filter, top-K. That's the lesson for the interview — "mergeable" is not a bonus property of sketches, it's the entry requirement for anything you want to shard. Twitter later swapped Storm for Heron (SIGMOD 2015), keeping the API while fixing debuggability and back-pressure at scale; the two-view shape survived the engine change.
Walk the mechanics — this is where hand-waving gets caught. A shard owns one Kafka partition and, for the current minute, two structures:
For each event: hash the key four ways, increment those four cells, read the estimate back as the minimum of them. If that beats the heap's smallest element, admit the key (or update it in place if it's already there) and evict the smallest. A handful of nanoseconds, no allocation, no growth.
A sketch's counters only go up. You can't remove an event from a count-min sketch — decrementing corrupts every key sharing those cells. So "the last hour" can't be one long-lived sketch that forgets.
The fix: never forget within a bucket, throw away whole buckets. Each shard keeps a ring of 60 per-minute sketches plus 60 heaps. Each minute the pointer advances and the bucket now 60 minutes old is zeroed and reused — no allocation, no GC churn. "Last hour" is the sum of the 60 live buckets. The day window is the same trick one level up: 24 hourly sketches, each the merge of its 60 minutes, so a day costs 24 more sketches rather than 1,440.
Answering "top 100 of the last hour" is then two steps. Candidate generation: union the keys from all 60 minute-heaps — 60 × 1000 entries, deduplicating to maybe 8,000 candidates. Re-scoring: query all 60 sketches for each candidate and add. That's about 2 million memory lookups, a few milliseconds. Sort, cut at K.
Re-scoring is not optional, and it's the step people skip. A key can sit at rank 400 in each of 60 minutes — never in any minute's top 1000 — and still be rank 5 for the hour. Heap membership tells you nothing about that; only summing the estimates does. The heaps generate candidates; the sketches produce scores.
Think of a rain gauge made of 24 separate buckets, one per hour. To know the last day's rainfall, add up the buckets. At the top of each hour you don't pour water back out of the day's total — you tip out the oldest bucket entirely and start refilling it. "The last 24 hours" is then slightly coarse at the edges, and that coarseness is the price of never having to un-count anything.
Everything so far is engineering. Here is the part that's actually hard, and the part I'd spend most of the interview on: the deliverable is an ordering, not a number, and every trick that makes the counting fit in memory perturbs that ordering. A count-min sketch overestimates and never underestimates — so a key's score can only be pushed up, never down. That means the error never demotes a true winner; it only lets impostors climb. And impostors only matter in one place: the boundary at rank K, where the gap between the last key in and the first key out is small. So the real design question is not "how accurate is my sketch?" It is "is the gap at rank K bigger than my error bound?" Answer that with arithmetic and the whole approximation becomes defensible; hand-wave it and you've shipped a leaderboard whose bottom half is noise.
Two knobs. Width w (counters per row) bounds how far the estimate can exceed the truth: roughly (e / w) × N, where N is the events in that sketch and e is 2.718. Depth d (rows) sets how often a key blows past the bound anyway — about 2% with 4 rows, 0.7% with 5. Rows are cheap; I'll take 4, with 5 one config line away.
My numbers: a shard's minute-sketch carries N = 18M events with w = 262,144.
overshoot ≈ 2.718 / 262,144 × 18,000,000 ≈ 187.
So a key's per-minute estimate on a shard can be ~187 too high; summed over 60 buckets, the hour is up to ~11,000 too high. Safe? Entirely depends on the boundary. For the global top 10, tenth place in an hour has on the order of a million views — 11,000 is 1% of it, and the gap between rank 10 and 11 at the head of a power law is far bigger than that. Safe. For the K = 1000 explore page, thousandth place might have 60,000 views with neighbours within a couple of percent, so 11,000 of slack shuffles ranks 900 to 1100 between refreshes. Not a stable ranking; fine as an unordered "thousand popular things" set.
Who does that hurt? Nobody looking at the top ten. The product manager who wants rank 950 to hold still — I'd tell them to treat the tail as a set, not a ranking. And anyone paid by rank: a creator bonus keyed to "top 1000 this week" must never read the sketch, it reads hourly_counts. That's the requirement from minute two doing its job.
Two moves before anything clever. Widen the rows where counts pile up: going from 262,144 counters per row to 2,097,152 costs 32 MB per sketch and cuts the overshoot 8x. Cheap. Keep a deeper heap than you serve: track 1000 per bucket to publish a top 100, so keys hovering near the boundary stay in the candidate pool instead of falling out and becoming invisible to re-scoring.
One optimization I'd deliberately skip: conservative update, incrementing only the cells that currently equal the minimum. It cuts error substantially, but the sketch stops being a clean linear structure — and linearity is exactly what makes merging across shards and regions provably correct. I'd rather spend 30 MB than lose the merge guarantee, and I'd say that out loud, because it shows which property I'm protecting.
Two levels of merge, and they're different — which trips people up.
Within a region, I partition by key hash. Every event for #monsoon lands on the same shard, so that shard's count is the whole count. The regional merge is then a plain K-way merge of 32 heaps, no sketch merging at all, because no key is split. That's the payoff for paying the routing cost.
The obvious objection: doesn't hashing by key create a hot partition when something goes viral? Chapter p2c06 says hot keys are the enemy. Here, unusually, they aren't, and the reason is precise. Work per event is constant and tiny: four increments, no state growth, no allocation. A key taking 1% of global traffic sends 100K events/sec to one shard — a few percent of one core — and pre-aggregation turns even that into a handful of large increments per second. What a hot key does not do here is the thing that makes hot keys dangerous elsewhere: it doesn't grow one partition's state, and it doesn't serialize on a lock or a row.
The pinch point is one level up, at the aggregator: single writer, all the fan-in. I'll take that in the failure section, because the honest answer isn't "it's fine," it's "it's stateless enough to rebuild."
Across regions, partitioning by key is impossible — you can't ship 10M events/sec across oceans to reach a key's owner shard. Each region ingests locally, so the same key is counted in every region independently. Now sketch merging earns its keep.
Count-min sketches merge by adding cells element-wise, provided every sketch shares width, depth, and hash functions. Get this right in the room, because the three sketches each merge differently and interviewers test it: HyperLogLog (the distinct-count estimator) by per-bucket maximum, Bloom filters by bitwise OR, count-min by sum — because a count-min cell holds occurrences, and occurrences in disjoint streams add up.
Then the subtle half: merging the heaps is not enough. Take the union of candidate keys from every region's heap and re-query the merged sketch for each. A hashtag ranked 3 in India and 400 in the US and Brazil never appears in those heaps, but its global total puts it in the top 10 — only union-then-re-query finds it.
The honest caveat, volunteered: candidate generation from top-1000 heaps is a heuristic, not a guarantee — a key sitting at rank 1001 in all thirty regions and top-10 globally slips through. Two defenses: keep K' generous (10x is cheap), and measure the miss rate against the batch path daily.
Rank by volume and the list shows the same five things every day: the biggest creators, the perennial hashtags, "weather." None of that is trending. Trending means rising unusually fast for itself.
A supermarket's bestseller list is milk, bread, eggs, forever — useless as news. What the manager wants on the front display is the umbrella that sold 200 today when it normally sells 5: small in absolute terms, enormous relative to itself. Same data, different question — one asks "what is biggest," the other "what changed."
So I compute a second, separate score. Each candidate key carries a baseline rate — an exponentially weighted moving average of its per-minute count, updated as buckets close — and I score the present against it:
score(key) = (rate_recent + a) / (baseline_rate + a)
with a smoothing constant a that stops division by tiny numbers from producing infinite excitement. Two guards:
I keep this deliberately simple, and I'd say so: anything fancier is a modelling project, and a defensible simple score with named failure cases beats a hand-waved fancy one.
In December 2010, with WikiLeaks dominating the news, Twitter users noticed that #wikileaks was missing from the trends list despite enormous, sustained tweet volume — and accusations of censorship followed. Twitter published a blog post explaining the mechanism rather than the politics: the algorithm surfaces topics that are spiking now, not topics that are merely large, so a subject discussed heavily and steadily for days can be huge and still not trend, while a sudden burst about something small does. Twitter's help documentation has carried the same explanation ever since — trends identify what is popular now, rather than what has been popular for a while. That's the volume-versus-velocity split in public: two different metrics over the same counter, and users will assume you're running the one you aren't. The lesson is about expectations, not math — if your product says "trending," someone will eventually accuse you of rigging it, and your only defense is being able to explain the score precisely.
A trending list is a target, and a hundred bots hammering one hashtag a thousand times each look exactly like a real trend to a count-min sketch. The cheap fix: for velocity ranking, count distinct users, not events. Distinct counting is unaffordable for every key, but I only need it for the ~10,000 candidates clearing the volume floor — a HyperLogLog each (chapter p2c10) is 12 KB dense, so about 120 MB in total. Now a botnet must acquire distinct accounts to move the number, a far more expensive attack, and HLLs merge by per-bucket max, so it rides the same machinery.
I'd raise these before being asked, because half of them are invisible in normal metrics.
The aggregator dies. Single writer, obvious weak point — but it holds almost no unique state, since every input still sits in the shards' sketch rings. A new aggregator is elected (with a fencing token, chapter p2c03, so the old one can't wake up and publish over it), asks the 32 shards for their summaries, and republishes within seconds. Readers keep serving the last cached blob meanwhile: stale, not wrong. The 60-second TTL means an outage longer than a minute surfaces as an empty leaderboard rather than a frozen one that lies indefinitely — visible degradation over silent staleness, deliberately.
A shard dies. Its partition is reassigned, and the replacement starts with an empty sketch that would silently undercount every key it owns. Fix: checkpoint the sketch ring every 30 seconds (it's a fixed-size byte array — the cheapest state you'll ever snapshot), then replay from the checkpointed Kafka offset. Thirty seconds at 300K/sec is 9M events, a few seconds of replay.
Late events. A watermark is the stream's declaration that it believes all events up to time T have arrived, so buckets before T can be closed. I'd set it two minutes behind; anything later is dropped by the fast path and counted only by batch. Emit a metric for how much gets dropped — a mobile client that starts buffering for an hour shows up there first.
The correctness alarm — fast-versus-batch divergence. The monitor I care most about and the one nobody builds. Daily, compare the batch job's exact top 1000 for each closed hour against what the fast path published at the time: overlap@K (how many of the exact top 100 were in the published top 100) and median relative error on those counts. Alert below 0.95 overlap or above 2% error. Nothing else in the system can tell you the sketch has drifted — latency fine, CPU fine, no errors, list looks plausible, quietly wrong. Approximate systems need a correctness monitor precisely because approximation and corruption look identical from outside.
Also watched: publish lag (p99 under 8 seconds); consumer lag per partition; partition skew, the early warning for a hot key outrunning pre-aggregation; heap churn rate; cache hit ratio; and counter saturation, since a 4-byte counter caps near 4.3 billion.
Rollout. Sketch configuration affects correctness, so it ships in shadow: run the new width beside the old for a day, compare both against batch truth, cut over only when the new one wins. The serving blob is behind a flag naming its source, so rollback is a flag flip.
At 10x — 100M events/sec. The pleasing part: sketch memory doesn't move, because it never depended on key count. Two things break. The shuffle first — 10 GB/sec routed by key hash is the real bill, fixed by harder pre-aggregation at the collector plus more partitions. Then the subtle one: the error bound scales with volume, not keys. Overshoot is (e/w) × N, so 10x the events is 10x the absolute error and the rank-K boundary gets noisy while the memory graph stays flat. The fix is arithmetic — widen the rows 10x, 32 MB becomes 320 MB, still nothing. Knowing which resource is constant and which one silently degrades is the difference between having memorized a structure and having run one.
For years YouTube did something baffling: a video's public view counter raced to about 300 and then stuck — "301 views" — sometimes for hours, while comments and shares kept climbing. It became an internet meme, and the explanation YouTube published is a textbook version of this chapter's architecture. Below the threshold, views come from a fast, cheap counter that updates immediately. Above it, views are worth defending, so each goes through fraud and spam verification — real human, reloading bot, autoplay farm? — and while that slower audit ran, YouTube froze the public number rather than show a figure it might have to walk back. In 2015 they removed the freeze: counts update continuously now, with verification still running behind and adjusting numbers down when it finds inflation. A related choice came in 2019, when public subscriber counts became abbreviated, rounded figures while exact counts stayed in the creator's analytics. All of it is the decision this chapter forces on you: which number does the crowd see, which number is audited, and what happens in the gap? Twitter's answer was two merged views; YouTube's was first to freeze the display, then to show the estimate and correct it later. Both are defensible. Having no answer is not.
A passing senior answer sizes the exact hash map out loud before reaching for a sketch, commits to the two-path design with a stated reason, and knows exactly which population the approximation hurts. Probes to expect, close to verbatim:
| Decision | Why I committed | What it costs me |
|---|---|---|
| Count-min sketch + min-heap per bucket | Memory is constant in distinct keys; a spam wave can't inflate it | Overestimates only; ranks near the K boundary are unstable |
| Two paths (sketch fast, batch exact) | Display tolerates error; payouts and analytics do not, and replaying a lossy pipeline can't create exactness | Two codebases computing the same thing — the classic lambda tax, plus a divergence monitor to keep them honest |
| Kafka partitioned by key hash | Each key is complete on one shard, so the regional merge is a plain heap merge | A routing shuffle at 10M events/sec, and partition skew from viral keys (absorbed by pre-aggregation) |
| Ring of tumbling minute buckets | Sketch counters can't be decremented; discarding whole buckets is the only clean expiry | Window edges move in one-minute steps; 60 sketches per shard instead of one |
| Cross-region: add sketch cells, then re-query the union of heap keys | 32 MB of summary crosses the ocean instead of 600M events, and mid-ranked-everywhere keys survive | Candidate generation from top-1000 heaps is a heuristic — measured against batch, not proven |
| Velocity as a ratio to an EWMA baseline, with a volume floor | "Trending" means rising fast for itself, and the floor kills noise from tiny keys | Seasonality needs a week-old baseline; brand-new keys have none and need a cold-start rule |
| Distinct-user counting (HLL) for the top ~10K candidates only | Makes bot inflation expensive without paying for distinct counting on every key | ~120 MB of HLLs, and a second structure to operate and merge |
| 10-second freshness, published as one cached blob | 1M reads/sec never touch the pipeline; a 5-second edge TTL absorbs the hot key | The number on screen is up to ~10 s old — disclosed in the API rather than hidden |
Explain to a junior engineer, in five or six sentences, how a trending list can be correct enough to publish when the system never counts anything exactly.
Counting every hashtag exactly means one entry per distinct hashtag in memory — hundreds of gigabytes at our key counts, and a spammer can grow that on demand by inventing new keys. So instead we keep a fixed-size grid of counters where many keys share cells: a key's count lands in a few cells, and reading it back takes the smallest of them, which is never below the truth and sometimes above it. Beside the grid sits a small heap of the biggest keys, because the grid can score a key you hand it but can't list the big ones. The important part is that we're producing a ranking: the error only pushes keys up, so it can add an impostor near the bottom but can never knock a real winner out, and at the top the gaps dwarf the error. Anything needing real numbers — payouts, advertiser reports — reads an exact count from an hourly batch job over the raw log, late but right. And we compare the two daily, because a drifting approximation looks perfectly healthy in every other metric.
Sizing drill. Top-K over search queries: 8M searches/sec, and the query tail is brutal — call it 150M distinct queries per minute. Size (a) the exact per-minute hash map and (b) a sketch with 1,048,576 counters per row and 4 rows. Then compute the hour's overshoot on a 32-shard fleet and say whether K = 20 is safe.
(a) 480M searches a minute, 150M of them distinct, at ~90 bytes an entry with the string and map overhead: about 13 GB per minute-bucket, so ~800 GB for a sliding hour. Dead on arrival. (b) 4 × 1,048,576 × 4 bytes = 16 MB per bucket per shard; a 60-minute ring is under 1 GB per shard, and it doesn't move if distinct queries triple. (c) N per shard-minute = 480M/32 = 15M, so overshoot ≈ 2.718/1,048,576 × 15M ≈ 39, or ~2,300 across a 60-bucket hour. (d) The 20th most popular search of an hour gets millions of hits — 2,300 of slack is a rounding error, so K = 20 is safe by three orders of magnitude. Note the shape: exact fails on memory, the sketch's error is computed, and safety is judged against the gap at the boundary.
Now add per-country trending for 200 countries on top of global. The naive move is 200 more sketch rings per shard. Design something better and price it.
Don't multiply the sketches — multiply the keys. Keep one sketch per bucket — the widened 32 MB one from deep dive 2 — and hash the composite key country|key into it. Memory is unchanged (it never depended on key count); the cost is that N in the error bound is still total traffic, so every country inherits the same absolute overshoot. Then keep 200 small heaps: 200 × 1000 entries × ~50 bytes ≈ 10 MB. Total 42 MB per bucket, against 6.4 GB for 200 separate sketches. The failure case to name unprompted: for a country with a few thousand events a minute, the shared overshoot swamps its real counts — so the smallest countries get a dedicated narrow sketch (cheap, their N is tiny) or a batch-served list an hour behind. Same structure, two error regimes.
What breaks: your ingest team switches Kafka from key-hash partitioning to round-robin to "balance the load better." Nothing errors and the leaderboard still renders. What silently changed?
Every key is now smeared across all 32 shards, so no shard holds a complete count and the sound K-way heap merge breaks twice over. A key with 1/32 of its count on each shard may miss every local top-1000 and vanish from candidate generation — the classic distributed top-K failure, hitting broad-but-not-dominant keys hardest — and surviving candidates carry partial counts that don't rank against each other. The repair is the cross-region protocol: add the 32 sketches cell-wise, union the heaps, re-query, deepen K'. Cost: merging sketches instead of heaps, and a heuristic where you had a guarantee. Worth saying out loud — a sensible change by another team just degraded correctness with no error and no alert, and only the divergence monitor would catch it.
Requirement flip: advertisers now pay for placement in the trending module, so the published list must be exact and auditable. Everything else stays. What changes?
The sketch can't back a payable surface — no audit ends well with "our estimator overshoots by up to 11,000." But notice what actually made exactness unaffordable: unbounded keys. Paid placement applies to a registered set — say 100,000 campaigns — and an exact hash map of 100K counters is under 10 MB. So keep exact counters for the bounded set and the sketch for the open tail. I'd add idempotency keys on those events (chapter p2c02), because at-least-once delivery stops being a nuisance and becomes a correctness bug the moment counts pay out. The alternative — serve the paid surface from the batch table an hour behind, labelled as settled counts — is defensible too, but advertisers want their placement live.
"Design a system that counts ad clicks." The interviewer says it flatly, almost bored. You feel a flicker of déjà vu — didn't you just build a click counter in the Top-K chapter? Count things on a stream, sketch a count-min sketch, done in twenty minutes?
Then the interviewer adds one sentence, and the whole problem changes shape: "The counts are what we bill advertisers." Sit with that. In the trending chapter, if the #7 video was really #9, nobody on Earth noticed. Here, every count is a line item on an invoice. If you count a click twice, you charged a customer for something that didn't happen — at scale, that's not a bug, it's a lawsuit. If you drop clicks, your own company quietly loses revenue and never knows. Approximate answers, the entire toolkit that made Top-K easy, just became inadmissible on one of your two paths.
That's the game of this chapter: the same infinite stream of events, but with money-grade correctness on one side and real-time dashboards on the other — and those two requirements pull the architecture in opposite directions. The senior move is to notice, out loud, that they're different requirements and refuse to serve both from one machine.
Scope it out loud: "I'll build the pipeline from a click landing on our servers to (a) an advertiser dashboard and (b) a billing record. Serving the ad, ad auctions, and the ML fraud models themselves are out of scope — I'll leave a slot where fraud verdicts plug in. Fair?"
Functional requirements:
Non-functional, with committed numbers:
The throughput math first, because it's about to tell us something surprising. A click event — click ID, ad ID, timestamp, user context — is roughly 200 bytes. At the 50K/sec peak that's 10 MB/sec into the pipeline. Kafka does that on one modest broker. Raw storage: 1B/day × 200 bytes = 200 GB/day, call it 70 TB/year, and it compresses several-fold in object storage. Say the conclusion: "Throughput and storage are non-problems here. This is not a scale question wearing a disguise — it's a correctness question. My deep dives will be about guarantees, not throughput."
Now the estimate that actually sets the bar. Suppose the average cost-per-click is 50 cents. One billion clicks a day means roughly $500 million a day flowing through these counters. A 0.1% counting error — a rounding error by Top-K standards — is $500,000 a day, $180 million a year, either stolen from advertisers or leaked from your own revenue. That single line is why this chapter exists separately from the trending chapter: the acceptable error on the billing path is not "small," it's zero, and provable. Meanwhile the dashboard path genuinely doesn't need that — an advertiser watching a live graph cannot perceive 0.1%. Two bars, so two paths. The architecture writes itself from this paragraph.
In 2016 the Wall Street Journal revealed that Facebook had been overstating the "average time spent watching video ads" metric for about two years. The bug was a definition error in the aggregation: total watch time was divided only by the number of views that ran 3 seconds or longer, so short views vanished from the denominator while their seconds stayed in the numerator, and the average came out inflated — Facebook initially told advertisers by 60–80%, while the advertisers' lawsuit later alleged some figures were inflated far more. Facebook maintained the metric never fed billing and called the suit meritless, but in 2019 it still agreed to pay $40 million to settle with advertisers who argued they had bought video ads based on numbers that were wrong. Notice what the incident actually was: not an outage, not data loss — one aggregation formula, slightly wrong, running quietly at scale for two years. That's the failure shape of counting systems: they don't crash, they drift, and nobody sees it without an independent check. Hold that thought for the reconciliation section.
POST /v1/clicks — body {click_id, ad_id, event_ts, device_ctx}, returns 202 immediately. The click_id is a UUID minted by the client SDK at click time; a retry resends the same ID. That one field is the foundation of dedup (chapter 2.2's idempotency-key habit, applied at the front door).GET /v1/ads/{ad_id}/stats?from&to&granularity=minute|hour|day — dashboard reads, auth-scoped so an advertiser sees only their own ads.Three kinds of data, and keeping them separate is half the design: the raw click log (immutable, append-only, in object storage — the legal ground truth), minute aggregates (window_start, ad_id) → {valid_clicks, filtered_clicks, cost}, and billing/settlement rows derived from them. The minute is my billing atom: small enough that any dashboard granularity or invoice period is a sum of atoms, big enough that row counts stay sane. With 2M ads, most idle in most minutes, expect a few hundred million minute rows a day — a volume that matters when we pick stores.
Write path: the click hits the ingest service, which validates the shape, stamps an arrival time, and produces it to a Kafka topic — then returns 202. Nothing synchronous happens beyond "it's durably in Kafka." I'm partitioning the topic by ad_id, and this choice is load-bearing, so I'll justify it: hashing on ad ID means every event for a given ad lands on the same partition, which means one Flink task owns all of that ad's state — its window counts, its seen-click-IDs — with no shuffle, no reshipping events across the network to whichever machine holds their ad's counter. It also means a retried click (same click_id, same ad_id) lands on the same partition as the original, so dedup state can be local instead of global. The cost, and I'll name it before the interviewer does: a viral ad makes a hot partition. That's a deep dive.
The Flink job consumes, flags fraud (a cheap rule-based pass inline — datacenter IP lists, absurd per-device click velocity; the heavyweight ML scoring runs offline and adjusts at settlement), dedups by click_id, and counts into tumbling one-minute windows keyed by ad_id — tumbling meaning back-to-back and non-overlapping, so every click falls in exactly one bucket. Each closed window emits one row to two sinks: ClickHouse for dashboards — columnar, brutal at time-range aggregation, with materialized rollups minute→hour→day so a 30-day dashboard reads 30 rows, not 43,200 — and Postgres for billing. A separate dumb archiver copies the raw topic to S3 untouched.
Read path: the dashboard API hits ClickHouse rollups, filtered by the advertiser's own IDs (the table is ordered by advertiser_id, ad_id, window_start precisely because every real query starts with "my ads"). Billing reads Postgres — but only the settled rows, and what "settled" means is the second deep dive.
A busy shop runs two sets of numbers from the same receipt tape. On the wall there's a live sales counter — updated instantly, glanced at all day, and if it's off by a coffee or two, nobody cares. In the back, the accountant works from the receipt tape itself, slowly, and produces the books that the tax office sees. Same events, two consumers, two bars. The rookie mistake is making the wall counter accountant-grade (you'll rebuild the shop around a scoreboard) or letting the accountant copy from the wall counter (you'll go to jail). Keep the tape; feed both from it; never confuse which one you're reading.
Clock check: the loop is closed end to end. Now I announce the agenda instead of waiting for one: "The two problems this question exists to probe are how counts get into the billing store without ever double-counting across failures, and what happens to clicks that show up late. Delivery guarantees first."
This is the hardest part, and it's the payoff of chapter 2.9. Every component in this pipeline retries: the SDK retries the upload, the ingest producer retries the Kafka write, Flink replays from checkpoints after a crash, the sink retries failed batches. Retries mean duplicates, and a duplicate on this path is a customer charged twice. "Exactly-once" is the most oversold phrase in streaming — no network can deliver a message exactly once — so the real question the interviewer is asking is: through crashes and retries at every hop, how does the effect on the billing table happen exactly once? You must be able to walk a crash, step by step, and show where the duplicate dies.
Recap the machinery in three sentences (chapter 2.9 has the full story). Flink periodically checkpoints its state — every window's running counts, every operator's Kafka read position — as one consistent snapshot. On crash, it restores the last checkpoint and rewinds Kafka to the positions stored in that same snapshot, so state and input position always agree: internally, it's as if the failure never happened. The problem is the edge of the system: anything the job emitted between the last checkpoint and the crash has already left the building, and after restore it will be computed and emitted again.
Two honest ways to kill that duplicate at the billing sink:
UPSERT ... (window_start, ad_id) — an upsert being "insert this row, or overwrite the existing row with this key." Replayed emission, same key, same recomputed value — writing it twice changes nothing. This works here for a specific reason worth saying out loud: the natural key is stable and the value is deterministic. After restore, Flink recomputes the same windows from the same rewound input and produces the same rows. The duplicate doesn't sneak past the guard; it arrives and is absorbed.I'll commit to the idempotent upsert, and here's the reasoning. Exactly-once is an end-to-end property, and the cheapest place to buy it is the last hop — the same lesson as Stripe's idempotency keys in chapter 2.2. The upsert version has no coordinator, no transaction timeouts to tune, no wedged-checkpoint failure mode, and it works with any store that can upsert. The 2PC sink earns its complexity when the target can't upsert (an append-only ledger, a Kafka output topic where consumers must never observe uncommitted results) — that's a real case, it's just not this one. And there's a second net under the whole act, which is the next deep dive: the bill doesn't come straight from the stream anyway. The sink writes in checkpoint-aligned batches — a few tens of thousands of rows bulk-upserted every ten seconds, comfortably inside one well-kept Postgres — not row-at-a-time dribbles.
Now walk the crash, because the interviewer will ask: checkpoint 41 completes with Kafka offset 9,000,000 and the 12:04 windows mid-count. At 12:05:07 the job emits the 12:04 rows to both sinks. At 12:05:09 a task manager dies — after the upserts landed. Flink restores checkpoint 41, rewinds to offset 9,000,000, recounts 12:04 from the same events, emits the same rows again. Postgres upserts them: same key, same values, zero net effect. No double-billing, and I can say precisely why: replay determinism plus keyed idempotent writes. That two-clause sentence is the whole deep dive compressed; own it.
A user on the metro taps an ad at 09:12, event queued in the SDK; connectivity returns at 09:47 and the click arrives 35 minutes old. Which minute does it belong to? Its event time — 09:12, when it happened — not its processing time, when we saw it. Billing on processing time would make invoices depend on our queue lag and the user's subway schedule, which no advertiser would accept. So windows are event-time windows, and that creates the question that haunts all of chapter 2.9: how long do you hold the 09:12 window open, waiting for stragglers, before you dare emit a number?
The watermark is the pipeline's working answer: a moving claim that says "I believe I've now seen everything with event time before T," computed as the max event time observed minus a bound (I'll commit to 60 seconds). When the watermark passes the end of a window, the window fires and emits its count. For the stragglers behind even that, I'll allow 15 minutes of lateness: a late click inside that grace updates the already-emitted window, and the new, more complete row flows to both sinks — the upsert overwrites, which is exactly why the upsert design also solves late updates for free. Beyond 15 minutes, the click is diverted to a side output — a separate late-events stream, archived to S3. Not dropped. Never dropped. Just excluded from the real-time answer.
So the streaming numbers are honest but incomplete: they're missing whatever landed in the side output, plus whatever a bug miscounted. Which brings us to the sentence that resolves this whole chapter's tension, the one to say slowly at the whiteboard: the stream serves, the batch settles — and you bill from settled. Every night, a batch job (T+1: today it processes yesterday) reads the raw S3 archive — which contains everything, including every late arrival and every side-output refugee — dedups globally by click_id, applies the final fraud verdicts, recounts every (minute, ad) atom from scratch, and writes settled rows. The invoice is generated from settled rows only. The dashboard keeps serving streaming rows, clearly labeled provisional, because that's the path where two minutes of freshness matters and 0.1% doesn't.
The settle pass costs one job over 200 GB and buys three things: a hard ceiling on how wrong billing can ever be (one day, then truth reasserts itself), a second, independently computed opinion to compare against the stream — the drift detector Facebook's metrics bug lacked for two years — and the audit answer: every settled row is a deterministic recount of an immutable log we still have. When the two opinions diverge beyond 0.1% for any ad-day, a human gets paged before an advertiser does the math for us.
This is exactly how elections work. On election night you get live results — fast, updated every minute, accurate enough to call most races, and nobody treats them as final. Mail-in ballots postmarked on time (event time!) keep trickling in for days and are counted if they arrive within the legal grace period — arrivals after that aren't shredded, they're logged and handled by a separate process. Weeks later comes the certified count: slower, recounted from the physical ballots, and that is the one with legal force. Election night = your dashboard. Certification = your settlement job. A country that certified from the TV ticker would be insane; so is a billing system that invoices from the stream.
Duplicates arrive two ways: innocently (the SDK retried an upload that had actually succeeded; a user double-tapped) and maliciously (a bot farm hammering a competitor's ad to drain their budget). The innocent kind is solved by the click_id: the Flink job keeps, per ad, the set of click IDs it has already seen in its keyed RocksDB state — RocksDB being the on-disk key-value store Flink parks large state in, so the set never has to fit in RAM — with a 24-hour TTL that expires old IDs. A duplicate finds its ID present and is dropped before counting. Because we partitioned by ad_id, the retry provably lands on the same task as the original — the state lookup is local, no distributed set needed. A Bloom filter (chapter 2.10) could shrink that state, but its false positives would drop real clicks — a silent undercount on a money path — and RocksDB state is disk-backed and cheap, so exact wins here. And once more, defense in depth: the settle job dedups globally by click_id anyway, so anything that slips the streaming dedup dies at settlement.
The malicious kind is a different discipline — fraud detection is its own team at any ad company — so architect for it rather than solving it: an enrichment stage flags each click valid or filtered using cheap online signals (datacenter IP ranges, impossible click velocity per device, missing interaction signatures), and crucially it flags rather than drops. Filtered clicks flow through the whole pipeline and the raw log, excluded only from billable counts. That means the fraud model can be re-run retroactively at settlement with heavier offline ML, verdicts can be audited, and advertisers can be credited after the fact when detection improves. Never delete evidence in a system whose job is producing evidence.
In 2005, Lane's Gifts & Collectibles — an Arkansas retailer — led a class action claiming Google billed advertisers for fraudulent clicks. Google settled in 2006 for up to $90 million in advertising credits, but the technically interesting part is what the settlement required: an independent expert, NYU professor Alexander Tuzhilin, was brought in to examine Google's invalid-click detection from the inside. His report concluded Google's efforts were reasonable — and it publicly documented the architecture: layered filters, most fraudulent clicks discarded proactively by automated real-time filters before they ever reached a bill, backed by offline analysis and manual investigation, with credits issued retroactively for whatever got through. Two lessons for your whiteboard. First, "dedup and fraud filtering" isn't gold-plating — it's the difference between an invoice and a legal liability, and a court once made a professor audit exactly this pipeline. Second, notice the shape Tuzhilin described is the shape we just drew: cheap online filters inline, expensive offline judgment at settlement, money corrected after the fact. The industry converged there because billing wrong is worse than billing late.
A month in, you discover the dedup TTL had a bug and some region overcounted by 0.3% for nine days. The fix for the future is a deploy; the fix for the past is chapter 2.11's playbook, and this system was built for it. The raw log is immutable and complete, so: stand up the corrected job as a second reader over the S3 archive (Kafka's retention window will be long gone — that's why the archive exists), replay the nine days into shadow tables, diff shadow against live settled rows, eyeball the diffs until they're explainable, then swap in a verified transaction and issue credits for the deltas. Nothing is edited in place; the old wrong numbers remain on record next to the correction, because an audit trail that can be quietly rewritten isn't an audit trail. Announce this unprompted in the interview — reprocessing is the question interviewers hold in reserve, and answering it before it's asked is a senior tell.
Flink job crashes. Covered in the crux walk: restore checkpoint, rewind Kafka, recount, idempotent upserts absorb the re-emissions. The user-visible symptom is dashboards going stale for the recovery minutes — annoying, not costly. The metric that matters is watermark lag: how far event-time progress trails the wall clock. It's the single best health signal for the whole pipeline, because everything downstream — window firing, dashboard freshness — keys off the watermark. One subtle trap from chapter 2.9: the watermark is the minimum across partitions, so one idle partition (an ad topic partition with no traffic) can silently hold back every window in the job. Configure an idleness timeout so quiet partitions are excluded, and alert when watermark lag exceeds a couple of minutes.
The viral ad. Partitioning by ad_id put all of one ad's traffic on one partition — hot-partition roulette, straight from chapter 2.6. A Super Bowl campaign doing 50K clicks/sec alone would swamp a single Flink subtask while its neighbors idle. The standard fix, and I'd build it behind a hot-key detector rather than for every ad: salt the key — split ad_777 into ad_777#0…#15 spread across partitions, each subtask counting a partial sum, with a tiny second-stage task merging the 16 partials per minute into one row. Counting is the best possible case for salting because sums merge exactly — no top-K merge headaches, no sketch error. Cost: an extra hop and a merge operator, paid only by hot keys.
Sink trouble. If Postgres slows or a bulk upsert fails, the sink retries and backpressure propagates up the job — the slow sink makes the operator ahead of it slow, and so on back to the source, until the job simply reads Kafka more slowly. Kafka absorbs the backlog (that's what it's for; chapter 2.7's lesson that the queue is your shock absorber). Alert on consumer lag and checkpoint duration; a checkpoint that used to take 5 seconds taking 90 is your earliest warning that state or a sink is sick.
Silent wrongness. The scariest failure in a counting system is the one with no error log — a fraud rule misfiring, a dedup TTL bug, a timezone slip in windowing. Two independent detectors: the daily stream-vs-settled diff (already built), and canary clicks — synthetic clicks with known IDs injected continuously at the front door, then verified to appear exactly once in both sinks at the right minute. LinkedIn built exactly this as its Kafka audit system: count events per stage per time bucket, compare stage against stage, alarm on loss. End-to-end counting checks catch entire categories of bugs that per-component health checks can't see.
At 10x (10B clicks/day, 500K/sec peaks): Kafka and the archiver scale by partitions — so over-provision partition count on day one (say 256), because repartitioning later breaks the ad→partition stickiness the dedup design leans on. Flink scales with partitions; RocksDB state grows linearly and checkpoints must go incremental. ClickHouse ingest of a few thousand row-batches per second doesn't blink. The settle job is embarrassingly parallel over S3 files. Nothing structural changes — which is the mark of a design whose hard parts were guarantees, not throughput, all along.
What the interviewer probes next, nearly verbatim:
| Decision | Why | What it costs |
|---|---|---|
Partition Kafka by ad_id | Stateful aggregation and dedup with no shuffle; retries land where the state lives | Viral ads create hot partitions — needs detection + key salting with a merge stage |
At-least-once + idempotent upsert on (minute, ad), not a 2PC sink | Same effectively-once outcome with no coordinator, no wedged-transaction failure modes; late updates come free | Only works because the key is stable and recomputation is deterministic; an append-only sink would force 2PC back in |
| Event-time windows, 60s watermark, 15-min allowed lateness | Bills reflect when clicks happened, not our queue lag | Results are provisional and can revise for 15 minutes; watermark lag becomes a metric you must watch |
| Bill from T+1 batch settlement, not the stream | Complete input (all late events), global dedup, independent check on the stream, replayable proof | Invoices trail reality by a day; you run and own a second compute path |
| Tumbling 1-minute atoms | Any dashboard granularity or invoice period is a sum of atoms | Hundreds of millions of rows/day — forces rollups on the analytics (OLAP) side |
| Fraud flags, never deletes | Retroactive re-scoring, auditable verdicts, advertiser credits possible | Filtered traffic flows through and is stored — you pay to keep what you don't bill |
| Immutable raw log, 13 months | The audit answer and the backfill fuel; corrections are append-only | ~70 TB/year raw (much less compressed) — cheap, but someone owns lifecycle and access control |
A junior asks: "We have a real-time pipeline that's 99.9% accurate. Why do we bill from a slow overnight batch job instead?" Answer in five sentences.
Because 99.9% accurate on half a billion dollars a day is half a million dollars of wrong bills every day. The stream has to answer now, so it must guess about clicks that haven't arrived yet — phones come back online hours later — and a guess is fine for a dashboard but not for an invoice. The batch job runs a day later over the complete raw log, so it has every late click, can dedup globally, and computes the same number every time you run it — which also means we can re-run it in front of an auditor to prove a bill. It's also our alarm system: two independent computations of the same counts, compared daily, catch silent bugs that neither would catch alone. So the stream serves the dashboard, the batch settles the money — fast where fast matters, exact where exact matters.
Add budget pacing: an ad must stop being served once its daily budget is spent. What does this new consumer of your counts need, and why does neither of your existing paths serve it well?
Pacing needs a low-latency, always-available, approximately-right running spend per ad — seconds fresh, because at 1K clicks/sec a viral ad burns budget between dashboard refreshes. Settled data is a day late (useless); even the streaming sink trails by the window + watermark delay. So add a third, deliberately cheap consumer: a lightweight aggregator (or the Flink job itself) pushing running spend into Redis with sub-second freshness and no exactness guarantee, and stop serving at ~95% of budget to leave a safety margin for in-flight clicks. Overspend beyond what you can bill is eaten by the platform as over-delivery credit — which is exactly how real ad platforms handle it. The senior insight: a third consumer with a third freshness/accuracy contract gets a third path; don't torture the billing path into being fast or the dashboard path into being exact.
Super Bowl: one ad spikes to 100x its normal rate for 20 minutes while global traffic is 3x. List, in order, the first three things that degrade and your mitigation for each.
(1) The hot ad's Kafka partition and its Flink subtask saturate first — one key's traffic can't spread by adding machines. Mitigation: hot-key detection flips the ad to salted sub-keys (ad#0…#15) with a second-stage merge; sums merge exactly, so this is safe for money. (2) Checkpoint duration balloons as that subtask's state and in-flight backlog grow, delaying sink batches — alert on checkpoint time, use incremental checkpoints, and let Kafka absorb lag rather than shedding events (never shed on a money path). (3) Watermark lag rises as the backlog delays processing, so dashboards go stale for everyone — communicate staleness in the UI ("data delayed ~4 min") rather than pretending. Note what does not break: billing correctness. Backlog delays settlement inputs by minutes, and T+1 settlement doesn't care.
Requirement flip: the CFO demands that dashboards always exactly match what will be invoiced. What are your options, and what does each cost?
Option A: dashboards read only settled data — exact match achieved, but everything is a day stale, which kills the "how is my campaign doing right now" use case. Option B: show both numbers, labeled — "live (provisional)" and "billable (settled through yesterday)" — which is what mature ad platforms actually do; costs only UI honesty. Option C: make the stream invoice-grade in real time — and here you must say the hard truth: it's impossible without changing the definition of a click, because a click that hasn't arrived yet (offline device) cannot be in any real-time number. You'd have to bill on processing time, making invoices depend on queue lag — advertisers would rightly refuse. The senior answer is B plus educating the CFO that the mismatch isn't sloppiness; it's the physics of late-arriving data made visible.
An advertiser disputes their invoice for the 3rd of last month, claiming they were billed for ~2% more clicks than their own analytics show. Walk the audit path end to end.
Pull the settled rows for their ads for that day, then the raw events behind them from the S3 archive — every billed click with its click_id, timestamp, and fraud verdict. Re-run the settlement recount for that ad-day and show it reproduces the invoice exactly (determinism is the point of settling from an immutable log). Then reconcile definitions against their analytics: the usual suspects are their client-side tracker missing clicks that bounced before their page's JS loaded, our fraud filtering (we bill fewer than raw), and timezone or attribution-window mismatches. If a real discrepancy remains — say our dedup missed a duplicate class — the raw log proves it, you credit the delta, fix the job, and backfill via shadow tables. The design's whole audit story is one sentence: we can recompute any bill from evidence we never edit.
A dedup bug shipped Monday and was found Thursday: one SDK version's retries used fresh click_ids, so its duplicates were billed. Kafka retention is 7 days. Plan the correction.
First, stop the bleeding: hotfix the SDK signature server-side (fingerprint dedup on ad_id + device + timestamp bucket for that SDK version as a stopgap). Correction: you don't need Kafka at all — the raw S3 archive has the affected days (this is why the archive, not Kafka retention, is the system of record for replay). Write the corrected dedup as a batch job over Mon–Thu raw data into shadow settled tables, diff against live settled rows to size the damage per advertiser, review, then swap and append credit line items — never edit the original rows. Publish the incident to affected advertisers with the credits. Postscript for the interview: this is chapter 2.11's shadow-table migration applied to data instead of schema, and the reason "raw log retained 13 months" was a requirement, not a nice-to-have.
ad_id co-locates each ad's aggregation and dedup state so no shuffle or distributed set is needed — at the priced-in cost of hot partitions, fixed by key salting with an exact merge.(window, ad); the 2PC transactional sink is the honest alternative you reach for only when the sink can't upsert."Design a payment system." The interviewer says it flatly, like it's just another question in the gauntlet. It isn't. This is THE senior filter — the problem where candidates who've been coasting on scale vocabulary finally run out of road. There's no celebrity fanout here, no petabytes, no clever cache. There is only money, and one unforgiving rule: you may never charge someone twice, and you may never lose track of a cent.
Here's the moment the whole interview orbits. A user taps Pay. The spinner spins. Ten seconds pass. The request times out. The user taps Pay again. Question: did you just charge them once, twice, or zero times? If your design can't answer that question with certainty — not "probably once," certainty — nothing else you draw matters. Mid-level candidates hear "payment system" and start talking about throughput. Senior candidates hear it and start talking about retries, ledgers, and reconciliation, because they know the dirty secret: payments is a correctness interview wearing a scale interview's clothes.
This walkthrough plays it at the E5 bar: committed choices, numbers that gate decisions, and most of the time spent exactly where the fire is.
First sentence at the whiteboard: "I am not building card rails." Talking to Visa, vaulting card numbers, fighting fraud rings — that's what payment service providers (PSPs) like Stripe and Adyen exist for, with thousands of engineers and a decade of compliance behind them. I'm building the payments layer of a commerce platform — think a marketplace or food-delivery app — that sits on top of a PSP. Saying this out loud is itself a senior signal: knowing where your system ends is as important as knowing what's inside it.
Functional scope: charge a customer's card at checkout through the PSP; full and partial refunds; an internal ledger tracking every movement of money and each merchant's balance; receiving the PSP's webhook notifications; and lightweight merchant payouts (we compute what we owe; the PSP moves the money). Out of scope: fraud scoring, FX, subscriptions, and building our own card vault.
Non-functional, in priority order — and the order is the design:
Scale commitment: a big platform — 10 million payments a day, Black Friday peaks 10x normal hour. Watch what those numbers do next, because it's the opposite of what most candidates expect.
10 million payments a day is about 115 per second average. Peak, call it 10x: ~1,200/sec. For storage: a payment with all its audit events is a few KB — 10M/day × 5 KB ≈ 50 GB/day, under 20 TB a year, and it's append-mostly, partitioned by month. Now say the conclusion that filters seniors from mid-levels:
"This is small. Even a huge commerce platform does hundreds to low thousands of payments per second — one well-run Postgres handles the write load with room to spare. So I will spend zero minutes on sharding and all my deep-dive time on correctness: idempotency, the ledger, and reconciliation. This is a correctness interview, not a scale interview."
That paragraph does real work. It's honest (Stripe cleared over a trillion dollars in 2023 and over $1.4 trillion in 2024 — and that still averages out to only low thousands of payments per second across the entire planet). It commits the architecture: one CP relational database at the core — CP as in the CAP trade-off, meaning that if the network splits we refuse writes rather than serve two different versions of the money truth. And it announces the deep-dive agenda instead of waiting to be led. The estimate drove three decisions in thirty seconds — that's estimation doing its job.
POST /payment-intents with header Idempotency-Key: <client-generated-uuid> and body {order_id, amount_minor, currency, payment_method_token} → 201 {intent_id, state}.GET /payment-intents/{id} → current state. The order page polls this.POST /refunds (also idempotency-keyed) with {payment_intent_id, amount_minor}.POST /webhooks/psp — the PSP calls us: payment.succeeded, payment.failed, refund.succeeded, chargeback.opened (a chargeback is the cardholder's bank yanking money back out of your account after the customer disputes a charge)…That header carries the whole chapter. An idempotency key is a unique string the client attaches to a request so the server can recognise a repeat of that exact request and answer it without doing the work a second time. Send it once, send it ten times, one charge happens. Deep dive one is where it earns its keep.
Notice amount_minor: money is an integer count of the smallest currency unit — 4999 cents, never 49.99 dollars. Floating point cannot represent 0.1 exactly; sum a few million float payments and cents evaporate. Integers plus a currency code (whose exponent tells you yen have no decimals) is the only acceptable answer, and interviewers do check.
Core tables: payment_intents (id, order_id, amount_minor, currency, state, psp_ref, attempt, version, timestamps), idempotency_keys (key, request_hash, response_code, response_body, intent_id), payment_events (append-only audit: every state transition, who/what/when), ledger_entries (append-only, the money truth — deep dive two), webhook_events (psp_event_id unique — dedupe), refunds, payouts. The token in payment_method_token matters for compliance, and we'll return to it — but first, the machine.
Write path: checkout posts an intent → in one database transaction we insert the intent row (state created), the idempotency-key record, and an audit event → a worker claims the row, transitions it to processing, and calls the PSP's charge API with its own idempotency key → the PSP's answer (and later its webhook) drives the state to succeeded or failed, writing ledger entries on success. In practice the API kicks the worker immediately after commit so the happy path feels synchronous — the poller is the same code path running as a safety net, sweeping up anything a crash left behind.
Read path: the order status page polls GET /payment-intents/{id} against a read replica. Consistency stance, committed: the ledger and intent store are CP — one writer, synchronous replica, and if we must fail over, we accept a short write pause rather than ever accepting two versions of the money truth. Product surfaces — order page, merchant dashboard — are eventually consistent read replicas; a status arriving two seconds late is invisible, a forked ledger is unforgivable.
Every intent lives on a strict state machine, and this is where auditability starts:
Minute twenty. The loop is closed. "Three deep dives, hardest first: surviving retries at the PSP boundary, then the ledger, then reconciliation."
Your database and the PSP are two computers that cannot share a transaction. Every payment is therefore a dual-write: a row in your Postgres and a charge on Stripe's side, with a network in between that can fail after the money moved but before you heard about it. Retries — from users double-tapping, from mobile networks, from your own timeouts, from the PSP's webhook redelivery — hammer this seam constantly. Everything distinctive about payment architecture — idempotency keys, intent-first writes, the poller, reconciliation — exists because of this one seam. The candidate who treats "PSP call + local write" as one atomic step has already failed; the interviewer is watching for whether you see the gap and build for it.
Walk the horror scenario first, because it justifies everything after. You call the PSP to charge $49.99. Ten seconds, no response, timeout. You are now in one of three worlds and cannot tell which: (1) the request never arrived — no charge; (2) the request arrived, the card was charged, and the response died on the way back; (3) the PSP is still processing. Now the fatal move: a "helpful" blind retry. In world 2, that's a second $49.99 charge. Customer support tickets, chargebacks, a Reddit thread with your app's name in it. The rule, tattooed: on timeout, never blindly retry a payment — query. Ask the PSP "what happened to the charge with this key?" and act on the answer. Ambiguity is resolved by asking, never by re-doing.
Now build the machinery that makes retries safe end to end — the full mechanics from chapter 2.2, deployed for real:
(key, hash(request_body)) and the intent row atomically. On a retry: same key, same hash → replay the stored response, byte for byte, without re-executing anything. Same key, different hash → 409, because the client is confused and guessing is how money gets lost. The atomicity matters: if the key record and the intent could commit separately, a crash between them would leave a key with no intent (retry does nothing forever) or an intent with no key (retry creates a sibling charge).202 — in flight, poll here. At no point do two executions run.psp_key = intent_id + ":" + attempt. Every retry of attempt 1 sends the same PSP key, so even if your process crashes mid-call and a sweeper re-sends, the PSP deduplicates on its side. A genuinely new attempt (card declined, user typed a new card) increments attempt — a new charge on purpose, not by accident. Two layers of idempotency, yours and theirs, seam covered from both sides.Paying through a PSP is sending rent via a courier. You seal cash in an envelope with a serial number, log the number in your diary, and send the courier. The courier doesn't return. Did the landlord get paid? The catastrophic move is sealing more cash in a fresh envelope. The correct move: send a courier to ask — "did envelope #4471 arrive?" And because the landlord also logs envelope numbers and refuses to accept the same number twice, even a duplicate courier can't double-pay. Your diary is the intent row; the serial number is the idempotency key; the landlord's log is the PSP's dedupe. The analogy breaks in one place: real couriers eventually show up or don't — networks can stay ambiguous for minutes, which is why the diary entry must be written before the courier leaves.
Which brings us to the dual-write trap itself. Why not wrap the PSP call inside the database transaction? Because an HTTP call is not rollback-able — if the transaction aborts after the PSP said yes, the money moved and your database remembers nothing. (Holding a transaction open across a ten-second network call is also how you starve your connection pool.) The fix is the outbox pattern from chapter 2.1, in its payments costume: the intent row is the outbox message. State created means "a PSP call is owed." The worker claims it, marks processing, calls out with the deterministic key, records the outcome. Crash before the call: the row still says created; the poller re-claims it. Crash after the call but before recording: the row sits in processing past its deadline; the sweeper queries the PSP by K:1 and writes what actually happened. Every crash window has a recovery story, and every recovery is a query or an idempotent re-send — never a blind re-charge.
Last piece of this dive: webhooks in. The PSP confirms outcomes by calling you — and their delivery contract is "at least once, in any order." You will receive duplicates; you will receive a refund.succeeded before the payment.succeeded it depends on. So: verify the signature, insert the PSP's event id into webhook_events with a unique constraint (duplicate → ack and drop), then apply the event through the state machine with a version check — compare-and-set on the intent's version column. If the transition is illegal from the current state, don't force it: park the event and retry shortly, or refetch the object's current truth from the PSP's API. The state machine is the bouncer; events don't get to skip the line.
The definitive public treatment of this design is Stripe's own. In 2017, Brandur Leach of Stripe's API team published "Designing robust and predictable APIs with idempotency," documenting how Stripe handles exactly the ambiguity above at planetary scale: every mutating request accepts an Idempotency-Key header; the server saves the result of the first execution — status code and body — and any retry with the same key gets that stored response replayed instead of a re-execution, whether the first attempt succeeded or failed. Keys are kept for 24 hours, long enough to outlive any realistic retry storm, then recycled. Stripe's client libraries even generate keys and retry automatically with exponential backoff, because Stripe would rather design for the network's failure than hope around it. The detail worth quoting in an interview: Stripe replays the response even for errors, because a retried request must be indistinguishable from its original — that's what makes "just retry it" a safe instruction instead of a dangerous one. When one of the world's largest payment companies documents its survival kit, and item one is idempotency keys, that tells you where the fire is.
Where does the money truth live? Not in a balance column that gets UPDATE-ed. In a double-entry, append-only ledger — the design accountants have stress-tested for five hundred years, for the same reason we want it: fraud and error detection by construction.
Double-entry: every movement of money is recorded twice — it leaves one account and enters another, a debit and a credit, and the amounts of every transaction sum to zero. Money is never created or destroyed in the ledger; it only moves between named accounts: psp_receivable (the PSP holds money it owes us), merchant_payable (we owe the merchant), platform_revenue (our cut), refunds_clearing. A $50.00 sale with a $5.00 platform fee:
Append-only is non-negotiable: rows are inserted, never updated, never deleted. A mistake is corrected by appending a reversing transaction that visibly undoes it — the error and its correction both stay in the record forever. This is what makes the ledger auditable: an UPDATE can silently rewrite history; an append-only log cannot lie about the past, only add to it. Database permissions enforce it — the application role literally has no UPDATE grant on ledger_entries.
Balances are derived, never authoritative. A merchant's balance is SUM(entries) over their account. Summing millions of rows per dashboard load is silly, so we keep a cached balance row, updated in the same transaction as each entry — but the cache is a view, and a nightly job recomputes every balance from raw entries and compares. A mismatch pages someone, because it means a code path moved money outside the ledger discipline. Cache for speed, recompute for truth, alert on divergence — the invariant check IS the design.
The ledger is a row of labeled jars and a rule: you may only move beads between jars, never pocket one or conjure one, and every move gets a line in a notebook — "3 beads, jar A → jar B." Count all the beads any night and the total must match the notebook exactly; if it doesn't, someone reached into a jar without writing it down, and the notebook tells you exactly when the count last held. A single "balance" column is just a sticky note on one jar saying "about 40 beads" — when it's wrong, there is nothing to check it against. The notebook is why double-entry survived five centuries: it isn't a record of balances, it's a record of movements, and movements can be replayed and audited. Where the analogy breaks: our beads also sit in other people's jars (the PSP's), which is why the next section exists.
Everything so far assumes our view and the PSP's view agree. They will drift. Webhooks get lost past their retry horizon; a sweeper hits an edge case; someone refunds a charge in the PSP's web dashboard, bypassing our API entirely; and — as the war story below proves — sometimes the card rails themselves misfire. Reconciliation catches what webhooks miss. It is not an afterthought; for a payments team it's a core product.
Every day the PSP publishes a settlement file: every charge, refund, fee, and chargeback they processed for us, with references and amounts. A batch job matches it line by line against our ledger on psp_ref, producing three buckets:
The unmatched-transaction workflow is part of the design, not an ops afterthought: breaks auto-retry matching for the lag window, then land in a review queue with the full audit trail attached (this is where payment_events pays for itself), and the reconciliation break count is a paging metric — nonzero after the lag window means an engineer looks today. A payments system that reconciles daily discovers its bugs in 24 hours. One that doesn't discovers them at tax time, or on Reddit.
In February 2018, Coinbase customers started reporting something terrifying: crypto purchases from weeks earlier were reappearing on their card statements as fresh charges — some users reported duplicates stacking up dozens deep, with bank accounts drained and overdraft fees piling on. The internet's instant verdict was "Coinbase double-charged us." The truth was stranger: the card networks had reclassified digital-currency purchases under a new merchant category code (to treat them like cash advances), and during that reprocessing, earlier transactions were reversed and re-charged down on the rails, below anything Coinbase's own code touched. It took days of joint investigation before Visa and the payment processor Worldpay issued a joint statement taking responsibility and confirming the erroneous charges would be reversed — explicitly stating the issue was not caused by Coinbase. Two lessons for your whiteboard. First: your code can be perfect and duplicates can still appear, because the rails beneath your PSP are themselves distributed systems that retry — which is exactly why reconciliation against settlement data is mandatory, not paranoid. Second: when money looks wrong, the company that can produce a complete, timestamped ledger of what it actually did clears its name in days instead of months.
Refunds reuse the whole machine: idempotency-keyed POST /refunds, a refund row linked to the original intent, its own small state machine, the PSP call driven by the worker, confirmation by webhook, reversing-direction ledger entries. Two rules with teeth: partial refunds are fine, but the running total of refunds must never exceed the original amount — enforced by taking a row lock on the payment while inserting the refund, so two concurrent $30 refunds on a $50 charge can't both pass the check. And refund-after-settlement is the normal case, not an edge case — the money already settled days ago, so the refund is a new movement in the next settlement file, matched by reconciliation like everything else.
Payouts (lite): a merchant's payable balance is already in the ledger; on a schedule we create a payout intent — same state machine, same idempotency, same outbox — ask the PSP to transfer, and append entries moving merchant_payable to payouts_in_transit and onward on confirmation. No new machinery; that's the payoff of building the spine well once.
PCI, one honest paragraph: raw card numbers never touch our servers, by construction. The checkout page embeds the PSP's hosted fields — an iframe the PSP serves — so the card number travels browser → PSP directly, the same shape as presigned-URL uploads from chapter 1.7 where clients ship blobs straight to storage and your servers only handle references. We receive a token, store the token, charge with the token. This collapses our PCI compliance scope from "fortress with quarterly audits" to a self-assessment questionnaire, and it means a breach of our database leaks no card numbers. Any design that routes raw card numbers through its own API servers is volunteering for the hardest compliance regime in commercial software to save one iframe.
PSP down for 30 minutes. Here's where intent-first architecture quietly pays off: checkouts keep committing intent rows; the worker's circuit breaker (chapter 2.7) opens and stops hammering; the queue of created rows just deepens. The product decision is honesty: for physical goods, tell the user "order placed — payment is processing, we'll confirm by email," then drain the queue when the PSP recovers. For instant digital delivery, fail fast and say so — pretending success you can't verify is how trust dies. What about failover to a second PSP? I'll commit: not in v1. Cards are tokenized per PSP — Stripe's token is meaningless to Adyen — so failover means storing the card twice, or moving to network tokens (card credentials issued by Visa and Mastercard themselves rather than by one processor, so they travel between processors), plus a second webhook dialect, a second settlement format, doubled reconciliation, and routing logic — a quarter's work and a permanent ops tax, purchased against the few hours a year a top-tier PSP is down. Queue-and-drain covers those hours. The trigger that flips this decision: PSP downtime materially breaching our checkout SLO more than once or twice a year, or cross-border fee arbitrage making multi-PSP routing pay for itself. Name the trigger, defer the complexity.
What we watch: auth success rate (a sudden drop is a PSP incident, an issuer outage, or our own bad deploy — it's the single most information-dense payments metric), webhook lag (PSP event timestamp vs our receipt), oldest intent stuck in processing (sweeper health), poller queue depth, reconciliation break count (page on nonzero past the lag window), and the ledger zero-sum check. And one alarm with a special status: the duplicate-charge detector — a continuous query for two succeeded PSP charges against one order. By construction it should never fire: the client key dedupes taps, the unique constraint dedupes concurrent requests, the deterministic PSP key dedupes crashed re-sends, and the PSP dedupes on its side. Four independent layers must all fail. We run the alarm anyway, over PSP-side data, precisely because it's built to be impossible — an alarm on an invariant is how you find out the day your invariant stops being true, and Coinbase's users can tell you the rails have their own opinions. If it ever fires, it pages like an outage and auto-opens a refund workflow.
10x growth? The unfashionable answer, again: 12K payments/sec is still Postgres territory with partitioning and a beefier box; the intent table is append-mostly and archives cleanly. What actually strains at 10x is the humans — a review queue growing linearly with volume is a slow-motion outage, so the metric that matters is the auto-heal rate: the fraction of reconciliation breaks resolved by machine. That's where the next engineer-year goes. Rollouts: state-machine code ships behind flags, canaried on a slice of traffic, with the old and new code required to agree on replayed historical events before the flag widens — you migrate a payments state machine the chapter 2.11 way, never big-bang.
The cautionary tale every payments engineer keeps in a drawer. In August 2012, market-maker Knight Capital deployed new trading code to seven of its eight production servers. The eighth kept the old build — and the deploy reused an old feature flag that, on that stale server, activated a dead code path called Power Peg, a long-retired test routine that bought high and sold low as fast as possible. At market open the eighth server started firing millions of unintended orders. There was no kill switch; while engineers debugged — at one point rolling back the good servers, making it worse — the system executed over 4 million trades in roughly 45 minutes, losing about $440 million, nearly four times the company's annual profit. Knight was effectively dead within the week, surviving only via emergency rescue financing, and the SEC's postmortem became required reading. It isn't a double-charge story — it's the deeper lesson under this whole chapter: systems that move money need staged rollouts that verify every node, kill switches that work in seconds, and independent controls that watch what the system is actually doing rather than what it's supposed to do. Reconciliation and invariant alarms are exactly that class of control, run daily instead of discovered during the fire.
Payment questions are probed almost entirely on the failure seams. Expect these, near-verbatim:
| Decision | Why | Cost accepted |
|---|---|---|
| Buy the rails: one PSP (Stripe/Adyen) | Card networks, vaulting, and fraud are a decade of someone else's scar tissue; our value is the layer above | Per-transaction fees; tokens locked to their vault; their outage is our degraded checkout |
| Single CP Postgres for intents + ledger | ~1.2K peak QPS needs no sharding; transactions give atomic intent+key+event writes | Vertical scaling ceiling; brief write pause on failover — accepted for one money truth |
| Client idempotency key; key+hash+response in one txn | Makes every retry — tap, timeout, crash — replay-safe end to end | Extra table and lock on hot path; clients must be taught to reuse keys on retry |
| Worker/outbox drives PSP calls from committed state | Every crash window resumable; no charge can exist that no row predicted | More moving parts than call-in-request; worker and sweeper to operate and monitor |
| Append-only double-entry ledger, derived balances | Zero-sum invariant catches bugs structurally; audit trail is the table itself | More rows, more storage; balance reads need caching plus a recompute-verify job |
| Daily reconciliation with paging on breaks | Catches what webhooks miss — including duplicates born on the rails themselves | Batch pipeline + human review queue to staff and continuously shrink |
| Queue-and-drain on PSP outage, no second PSP in v1 | Covers realistic downtime at near-zero complexity; honest UX | During outage: delayed confirmation, lost sales on instant-delivery goods — trigger named for revisit |
A junior asks: "Why can't we just call Stripe, and if it succeeds, save the payment to our database? One if-statement." Explain in five or six sentences why that's the most dangerous design in the codebase, and what we do instead.
Because the crash you didn't plan for lands exactly between those two steps: Stripe charges the card, then our process dies before saving — now a customer paid us and our database has no idea, which means no order, no refund path, and no way to even know it happened. The network makes it worse: if the Stripe call times out, we can't tell whether it charged or not, and retrying "to be safe" might charge twice. So we flip the order: first we save an intent row that says "we're about to charge $50 for order X," along with the client's idempotency key, in one transaction. Only then does a worker call Stripe, using a key derived from that intent, so even if we crash and call again, Stripe knows it's the same charge and won't run it twice. If anything dies mid-flight, a sweeper finds the unfinished row and asks Stripe what happened instead of guessing. The rule underneath: write down what you intend before you act, act idempotently, and resolve doubt by querying — never by re-doing.
Now add a stored-value wallet, with split tender: a $50 order paid $20 from wallet balance + $30 on card. Sketch what changes in the ledger and — the real question — what happens when the card portion fails after the wallet portion was taken.
The wallet is just another ledger account per user, so the happy path is two movements: wallet → merchant_payable ($20) and psp_receivable → merchant_payable ($30), each its own zero-sum transaction. The failure case makes this a saga (chapter 2.1): take the wallet money first as a hold (move to a wallet_holds account, not to the merchant), then attempt the card; on card success, commit the hold; on card failure, append a reversing entry releasing the hold. Never take wallet funds irrevocably before the card resolves, and never attempt the card first (you can't "hold" money on the customer's card and then discover their wallet balance changed — the wallet is the resource you control, so you reserve it first and make the external call the last risky step). Both the hold and the release are idempotency-keyed, because timeouts happen here too.
Black Friday: 10x traffic, and your PSP rate-limits you to 500 charges/sec while checkouts arrive at 1,200/sec. Design the behavior. What do users see, what do you monitor, and what breaks first if the surge lasts four hours?
The architecture already contains the answer: intents commit at 1,200/sec (Postgres doesn't blink), and the worker drains at the permitted 500/sec — the intent table becomes a durable queue, oldest first, with the circuit breaker respecting 429s. Users get the honest state: "order placed, payment processing," with confirmation following in minutes. Monitor queue depth and, more importantly, its derivative — at +700/sec net inflow, a four-hour surge queues ~10M intents, so the real risks are: authorization windows (a card auth attempted hours late fails more often — track success rate by queue age), inventory promised against payments that later fail, and email/notification storms on drain. First structural break: the product promise, not the database — which is why the senior answer includes pre-negotiating a higher rate limit with the PSP before the sale, the boring fix that beats any architecture.
The requirement flips: you now sell instant digital goods (game credits, delivered in-app within seconds). Which committed choices from the trade-off table does this overturn, and what replaces them?
Two flips. Queue-and-drain dies: you can't tell a user "credits processing, check back in an hour," so a PSP outage now means lost revenue per minute, and the multi-PSP trigger we named has effectively fired — the cost-benefit now favors a second PSP (or at least network tokens to make card credentials portable) despite doubled reconciliation. Second, delivery becomes part of the correctness story: granting credits must be idempotent against the payment intent id (webhook redelivery must not grant twice) and reconciliation gains a fourth bucket — payments settled but credits never granted — so you now reconcile money and fulfillment. The ledger, idempotency spine, and state machine survive unchanged; that's the sign they were the right spine.
A webhook arrives: refund.succeeded for intent p_777 — but your database shows p_777 still in processing. Walk through exactly what your system does, and name the two distinct root causes this pattern can have.
The handler dedupes the event id (new — proceed), then attempts the transition: processing → refunded is illegal, so it does NOT force the state. It parks the event for delayed retry and triggers a refetch of p_777 from the PSP's API. Root cause one: out-of-order delivery — the payment.succeeded webhook is in flight behind its sibling; the refetch (or the retried parked event, after the success event lands) resolves it within minutes. Root cause two, nastier: the success webhook was lost entirely AND the sweeper hasn't caught the stuck intent yet — the refetch heals it by walking the intent through succeeded-then-refunded with audit events for both. What you never do is jam the state to refunded directly: skipping succeeded means no settlement ledger entries were ever written, and now the refund reverses money the ledger never recorded receiving — an unbalanced ledger, found the hard way at reconciliation.
Your CFO asks: "If our duplicate-charge protections are 'structurally impossible' to break, why are we paying for a daily reconciliation pipeline and an alarm? Cut one." Give the senior answer in a few sentences.
Neither, and here's the argument: the four idempotency layers protect against failures in our request path — but "structurally impossible" is a claim about our code, and reconciliation is the instrument that continuously tests that claim against external reality. Coinbase's 2018 incident is the proof case: their code didn't double-charge anyone; the card rails reclassified and reprocessed over a million transactions, creating duplicates two layers below anything Coinbase controlled — only comparison against settlement data catches that class. An invariant without a detector is a belief; the alarm converts it into a monitored fact, and its cost is one query. You cut reconciliation the day you're willing to learn about missing money from customers instead of from a cron job.
"Design Google Docs." The interviewer writes two names on the whiteboard: Alice and Bob. "They both have the same document open. Alice types a word into the middle of a sentence. At the same moment, Bob deletes the start of that sentence. Neither has seen the other's edit yet. What does each screen show one second later?"
Feel the trap. If Alice's editor applies Bob's delete using the character positions Bob sent, the delete lands on the wrong characters — because Alice's insert already shifted everything over. Bob's screen has the mirror-image problem. Naive position-based edits don't just lag; they corrupt the document differently on each screen, and the two copies drift apart forever. This is not a scale problem — there might be exactly two users. It's a correctness problem, and it's the entire reason this question gets asked.
This is the payoff chapter for Part 2's collaborative-editing theory. Today you don't get to say "OT or CRDT, both are valid." At the senior bar you walk both on a concrete example, commit to one, and defend the commitment with this problem's actual requirements. Let's play it as the candidate.
"I'll scope to the collaborative text editor itself. Functional requirements: multiple people edit the same document in real time and every copy converges to the identical text; live cursors and presence (who's here, where they're typing); document history with named versions and restore; lightweight comments anchored to text ranges; sharing with view/comment/edit permissions, including revoking access while someone has the doc open; and tolerance for brief disconnects — a dropped Wi-Fi blip shouldn't lose typed characters. Out of scope: rich embedded objects, spreadsheets, and full offline-first editing — I'll say something honest about long offline later. Fair?"
Non-functional, with numbers I'll be held to:
"Let me get the shape of the load, because it decides the architecture. Say 100 million docs get opened per day, with 5 million open concurrently at peak. Average live editors per open doc: about 1.2 — most open docs have one active person; multi-editor docs are usually a meeting of 2–5; a 100-editor doc is rare enough to treat as a special case."
"Now the per-doc write rate. A fast typist produces about 5 characters per second, and clients batch keystrokes into an operation every 200 ms or so — call it at most 5 ops/sec per active typist, usually far less. Worst-case room: 100 editors × 5 ops/sec = 500 ops/sec into one doc, each fanned out to 99 others ≈ 50K small messages/sec. That's one busy but comfortable server core. A typical room does under ten messages a second."
"So here's the conclusion the numbers force: no single document is ever a throughput problem. The load is millions of tiny independent rooms. That means I can — and should — route all traffic for one doc to one server, because per-doc load fits on a fraction of a machine, and having one machine see every edit to a doc in one place is exactly what makes the merge problem tractable. The fleet-level problem is placement and failover of millions of small rooms, not scaling any one room. Fleet math: ~6M concurrent WebSockets at 30K per box is about 200 connection servers; an open doc's in-memory state is maybe 100 KB (text plus queues), so 30K rooms per box is ~3 GB — memory is fine."
Storage: an op is ~60 bytes; a heavily edited doc produces maybe 20K ops a day — about 1 MB. Keeping every op forever is cheap, and that will matter in a minute, because the op log is going to moonlight as the history feature.
REST for the boring parts: POST /docs, POST /docs/{id}/share, GET /docs/{id} (returns latest snapshot plus the op tail since it), GET /docs/{id}/history?from=seq. The interesting surface is the WebSocket session:
{op_id, base_seq, ops: [retain n | insert "text" | delete n]} — the standard span-based edit form: "keep 4 chars, insert this, keep the rest." Plus ephemeral cursor {pos} messages.{seq, op, author} broadcasts, ack {op_id, seq} for your own ops, and presence events.Two fields carry the whole design: base_seq — "the last server op I had seen when I made this edit" — tells the server exactly which concurrent ops to reconcile against, and op_id (a client-generated ID) makes resends after reconnect idempotent, straight from the idempotency chapter.
Data model: docs (metadata, ACL, head_seq), ops (doc_id, seq, op, author, client_op_id — append-only), snapshots (doc_id, seq, blob), comments (doc_id, anchor range, thread). Notice there is no "current text" column. The text is derived: snapshot plus ops since. The log is the truth.
"The write path: your editor applies your keystroke locally and instantly, batches it into an op, and sends it with its base_seq over the WebSocket to the doc's session server. The server reconciles it against any concurrent ops it has already accepted, appends it durably to the op log, assigns it the next sequence number, acks you, and broadcasts it to every other connected client, who reconcile it into their own local state. The read path: opening a doc fetches the latest snapshot plus the op tail after it, replays the tail, connects the WebSocket, and subscribes from the current sequence number."
Say the load-bearing sentence out loud: doc-sharding the connection layer means one server totally orders each document's edits. A total order — every op gets a global sequence number 1, 2, 3… within its doc — is the one gift that turns "distributed concurrent editing" into "a single queue plus some arithmetic." We can afford this gift because estimation told us no doc outgrows one server. Systems that can't centralize (peer-to-peer, offline-first) don't get the gift, and you'll see in a moment how much machinery they must buy instead.
Two people edit the same text at the same time; every screen must converge to the same result, and each user's own keystrokes must appear instantly, before the server has seen them. Everything else in this design — sessions, logs, snapshots — is solid Part 1–2 material. This is the part the question was invented to probe, and the interviewer wants three things: a concrete demonstration that you understand why naive merging fails, a working knowledge of both fixes (transform the operations, or restructure the data so merging is trivial), and a committed, defended choice between them for this product. Hedging here caps the interview.
Make it concrete. The doc says the cat. Both Alice and Bob are at version 0.
"fat " at position 4 → her screen: the fat cat."the ") → his screen: cat.Now exchange the ops naively. Alice applies "delete 4 chars at 0" to the fat cat → fat cat. Fine. Bob applies "insert fat at 4" to cat → cat fat… wait — position 4 of a 3-character string. Best case garbage, worst case a crash, and either way Alice sees fat cat while Bob sees something else. The positions in an op are only meaningful against the exact document state it was made on. Concurrency broke that assumption. Two families of fixes exist.
OT keeps ops position-based but transforms them before applying: given two ops made concurrently against the same state, a transform function rewrites each so it applies correctly after the other. Here: Alice's insert at 4 transformed against Bob's delete-4-at-0 becomes insert at 0 (her position slides left by the 4 deleted chars). Bob's delete at 0 transformed against Alice's insert at 4 is unchanged (his range sits entirely before her insertion point). Both screens reach fat cat. Draw it as a square:
A magazine's copy desk. Two proofreaders mail in corrections referencing line numbers — "fix the typo on line 12." But the first correction to arrive added a paragraph, so by the time the second letter is opened, its line numbers point at the wrong lines. The copy editor doesn't reject the late letter; she adjusts its line numbers to account for the edits already applied, then applies it. That adjustment is the transform, and the single copy desk everyone mails is the serialization point. The analogy's honest limit: real OT must also handle both letters editing the same line, where "adjust the numbers" needs genuinely careful rules — that's where OT's difficulty lives.
Here's the part candidates miss: OT's difficulty depends brutally on the topology. With no central server, any pair of concurrent ops anywhere may need transforming against any other, transforms compose along different paths, and correctness proofs get so hairy that several published peer-to-peer OT algorithms were later shown to be wrong. With a central server assigning a total order, the problem collapses: the server only ever transforms an incoming op against the ops it accepted after that op's base_seq — a simple loop — and each client only transforms server broadcasts against its own few unacknowledged ops. This server-ordered scheme is the Jupiter design out of Xerox PARC (1995), and it's the lineage Google Docs itself sits in. One server per doc isn't a scaling compromise. It's the move that makes OT small.
A CRDT (conflict-free replicated data type, from Part 2) attacks the root cause instead: positions were the fragile thing, so eliminate positions. Give every character a permanent unique ID at birth — (author, counter) — so the cat becomes t₁h₂e₃␣₄c₅a₆t₇. Alice's edit is now "insert my new characters after the character with ID ␣₄." Bob's edit is "mark t₁h₂e₃␣₄ as deleted" — deleted characters become tombstones, hidden but kept, so other people's edits can still anchor to them. Apply these two edits in either order on any replica: the tombstoned chars vanish from display, Alice's chars sit after ␣₄'s tombstone, everyone renders fat cat. No transform function at all — merging is just taking the union of edits, and the data structure's rules make the result order-independent. That's why CRDTs shine exactly where OT struggles: no server, peers syncing pairwise, weeks offline. The bill: every character carries ID metadata forever, tombstones accumulate, and the document is no longer a plain string but a specialized structure you must maintain, garbage-collect, and snapshot.
"For this product I commit to server-ordered OT, Jupiter-style. My reasons, tied to our requirements: (1) We already require a central server — permissions with mid-session revocation, durable history, and comments all need one, so OT's precondition costs us nothing extra. (2) Estimation showed every doc fits on one server, so the total order is free, and with it OT needs exactly one transform-pair for span-based text ops — a small, testable core. (3) The document stays a plain string: snapshots are just text, history is readable, memory per room is tiny — no per-character metadata or tombstone growth on billion-doc storage. (4) Our offline requirement is 'tolerate blips,' not 'edit on a plane for a week' — and blips are handled by the same transform machinery, as I'll show. I'd flip to a CRDT if the requirements flipped: offline-first, peer-to-peer, or end-to-end encryption, where no server may read — let alone transform — anyone's ops. The cost I accept: transform code must be exactly correct and deterministic everywhere it runs, client and server, forever — I'll treat it like crypto code: tiny, property-tested, versioned."
Chapter 2.4 told you what it costs to build OT: Google Wave burned two years on its transform engine, and the engineer who wrote it went on record that implementing OT sucks. Etherpad tells you the other half — what it costs to own one for a decade, which is precisely what the commitment above signs you up for. Etherpad launched in November 2008, the first web app of its kind to pull off genuinely real-time collaborative editing in a browser, and its engine is a changeset library called Easysync doing exactly what we just designed: the server keeps an ordered list of revision records numbered 0, 1, 2…, and a changeset that arrives built against an older revision gets pushed forward through every revision since. Google bought the company in December 2009 to fold the team into Wave, open-sourced the code that same month under Apache 2.0, and shut the original service down the following May. Then comes the part worth your attention. The community rebuilt the product from nothing — the original was Java and Scala, "Etherpad Lite" is JavaScript on Node.js, described by its own maintainers as an almost complete rewrite — and carried Easysync across unchanged. New language, new runtime, new people, same transform library, because the transform layer is the one piece nobody volunteers to re-derive.
Two details from that spec are what "treat transform code like crypto code" means in practice. First, Easysync's transform function is named follow, and the spec states it as an algebraic law rather than a behaviour: apply A then follow(A,B), or apply B then follow(B,A), and you must land on the identical document. That's the transform square from the diagram above, written as a property you can throw ten million random op pairs at overnight — which is how you find the delete-inside-a-delete case you forgot, instead of finding it inside a customer's dissertation. Second, the changeset format bans non-canonical encodings: two adjacent keeps that could be merged into one are illegal, and within a run of edits the deletions must be written before the insertions. That rule buys nothing at runtime. It exists so the same edit always serialises to the same bytes — which is what makes state hashes comparable across two implementations, and is what the rollout check later in this chapter (replay a slice of the log on both versions, compare hashes) quietly depends on.
Figma faced the same fork and chose the other path — for reasons that make both choices look right. In the 2019 engineering post "How Figma's multiplayer technology works," co-founder Evan Wallace explained that Figma's multiplayer isn't text: a design file is a tree of objects (frames, shapes, text layers) with properties. They looked at OT and judged it more complexity than their problem deserved. Instead they built something CRDT-inspired: each object property is a last-writer-wins register (two people set the same rectangle's fill concurrently → latest write wins, which is perfectly acceptable for a color in a way it never would be for prose), and child ordering uses fractional indexing. Crucially, Wallace was explicit that Figma doesn't run textbook decentralized CRDTs: every document still lives on one server process that orders changes, which lets them discard much of the metadata a true peer-to-peer CRDT must carry. So even Figma kept the one-server-per-doc serialization point. The lesson to say out loud: pick merge machinery by data shape and topology. Character-level intent in linear text → transforms. Independent object properties where last-writer-wins is fine → CRDT-style registers. Neither company chose by fashion.
Local echo in 16 ms means the client applies your edit before the server has seen it — optimistic apply. But then the client's document is ahead of the server's truth, and remote ops arriving from the server were built against a state your screen no longer shows. Every client therefore tracks its position precisely with a two-queue model. Learn it gently — it's four ideas:
base_seq 41) and not yet had acked. We allow only one in flight at a time — Docs-style — which keeps every case simple.Three events, three rules. A keystroke: apply to the screen now, append to pending. An ack for your in-flight op: it's now part of server truth; advance the base, promote the pending buffer to in-flight, send it. A remote op: it was made against server history, but your screen also contains your in-flight and pending edits — so transform the remote op over both queues before applying it to the screen, and symmetrically transform both queues over the remote op (your in-flight op's ultimate landing position just moved). Same transform function as deep dive 1, reused. Meanwhile the server does the mirror job: an op arriving with base_seq 41 when the server is at 45 gets transformed over ops 42–45, then appended as 46, acked, and broadcast.
Your chequebook versus your bank statement. The balance you act on is the bank's confirmed balance, plus the cheque you've mailed but the bank hasn't cleared (in-flight), plus cheques written and still in your bag (pending). When the statement arrives showing someone else's charge, you don't panic — you reconcile: adjust your running figure for the new entry while keeping your own uncashed cheques in the picture. The client does this reconciliation on every remote op, automatically, in a millisecond. Where the analogy breaks: the bank never reorders your cheques' meaning, but a remote edit really can shift where your pending edit lands — that's the transform step, and the chequebook has no equivalent.
And here's the cute detail interviewers love: cursors are just positions, so they transform too. When Bob's delete-4-at-0 arrives, Alice's cursor at position 8 slides to 4 by the same arithmetic that moved her ops. Presence and cursors ride a separate ephemeral message type on the same socket — never persisted, dropped on disconnect, rebroadcast on a timer — because nobody needs a durable record of where your cursor was on Tuesday. Comment anchors are the durable cousins: a comment on characters 120–134 is a stored range whose endpoints the server transforms as each op applies, so comments stay glued to their text as it moves. One transform function, four customers: ops, queues, cursors, anchors. Say that sentence in the interview.
Permissions get checked per-op, not per-connection. Checking only at WebSocket connect means someone whose edit access was revoked mid-session keeps writing until they disconnect — days, maybe. So the session server holds the doc's ACL in memory, an ACL-change event invalidates it, every incoming op is checked against it, and revocation also proactively closes the socket. The per-op check is the guarantee; the socket close is just courtesy.
Every accepted op is appended to a durable, replicated op log before the ack and the broadcast. That ordering is a promise worth stating: nothing any collaborator has ever seen can be lost. If the session server dies a millisecond later, the op is in the log; recovery replays it. Ack-before-durable would be the lie that loses someone's paragraph.
Replaying a year of ops to open a doc would be absurd, so every N ops (or on idle) the server writes a snapshot — the full text at sequence S. Opening a doc = latest snapshot + replay the tail. The cadence N is a knob: snapshot too often and you're writing the whole document over and over (write amplification — a 1 MB doc snapshotted every 10 ops writes 100 KB per keystroke-batch); too rarely and opening a doc replays a long tail (slow opens, slow recovery). N = 1000 ops bounds replay to a few milliseconds of CPU while keeping snapshot writes rare. Committed, and trivially tunable per doc size later.
Now the feature you get free: the op log IS the history feature. "See version history" replays the log to any point. A named version — "Final draft" — is nothing but a tag on a sequence number. Per-author attribution and those colored playback views come from the author field already on every op. Restore doesn't rewind the log (append-only means append-only); it computes the diff back to the tagged state and applies it as new ops, so restores are themselves visible, attributable history. In most systems, history is an expensive bolted-on feature; here it fell out of the write path. That's the sign the write path was designed right.
Offline, honestly: a brief disconnect is a non-event — the two-queue model already holds unacked ops, so the client reconnects, says "I'm at seq 2290," receives ops 2291–2317, transforms its held queue over them (the same rebase-by-transform), and resends with idempotent op_ids so a retry can't double-apply. That handles the metro tunnel. It degrades with absence: after a weekend offline, you're transforming hundreds of ops against thousands, and the merged result — while convergent — may no longer say anything anyone meant. The honest product answer is bounded offline: past a threshold, stop silently merging and surface a conflict UX ("your offline copy vs. current — review the differences"). Silent convergence to semantic nonsense is worse than an honest fork. If the requirement were true offline-first, that's my stated trigger to reopen the CRDT decision, not to stretch OT past its lane.
Session server dies. Thirty thousand rooms go dark. Clients detect the dropped socket, re-ask the directory, and land on a new server, which loads each doc's snapshot plus tail from the durable log — this is why durable-before-broadcast was non-negotiable — and resumes assigning sequence numbers exactly where the log ends. Each client sends its last-seen seq (gets the gap replayed) and resends unacked ops (deduped by op_id). Total user experience: a two-second "reconnecting" toast. Nobody loses a character.
Split-brain: two servers think they own doc 7. The directory failed over while the old server was slow, not dead — the classic zombie from the fencing chapter. If both append to doc 7's log, sequence numbers collide and clients see divergent truth: the one catastrophe this design must make impossible. So the directory issues a monotonically increasing session epoch per doc with each ownership grant, every log append carries the epoch, and the log rejects appends from any epoch older than the highest it has seen — a fencing token, exactly p2c03's medicine. The zombie's writes bounce off the log; it discovers it's dead and drops its sockets. The invariant, stated for the interviewer: correctness never depends on the directory being right, only on the log's epoch check being enforced.
What I watch (the on-call dashboard): op-apply p99 — full keystroke-batch-to-broadcast latency, the product's heartbeat; per-room transform queue depth — a doc whose incoming ops outrun processing is my hot-room early warning; reconnect success rate and time-to-resume after deploys and server loss; snapshot lag (max tail length) — a doc 50K ops past its snapshot means the snapshotter broke and recovery time is quietly growing; and WebSocket churn as a proxy for network or LB trouble.
Rollouts have a trap specific to this system: transform logic must produce identical results wherever it runs — old clients, new clients, the server. Ship a subtly changed transform to half the fleet and the same op pair reconciles differently on two screens: divergence, the unforgivable bug. So ops carry a schema version, servers hold transforms for N versions back, changes canary by doc cohort (a doc's whole room shares one code path — never mixed within a session), and a background checker replays samples of the log comparing state hashes across versions before anything ramps.
What breaks at 10x? Ten times the docs is the boring axis — 2,000 session servers instead of 200; doc-sharding scales linearly and the directory barely notices. The interesting 10x is audience per doc: a company all-hands doc with 20 editors and 10,000 viewers. Viewers don't need OT — they hold no pending edits, so they need only the ordered broadcast stream. I'd keep editors on the session server and hang viewers off a read-only fan-out tier of relays that tail the doc's op stream — the same fan-out tree the chat chapter built — so a viral doc never competes with its own editors for the serialization point.
This problem is a depth probe wearing a product costume. Interviewers report the same follow-up ladder almost verbatim:
| Committed choice | Why | Cost accepted |
|---|---|---|
| Server-ordered OT (Jupiter lineage) | Central server already required; total order per doc is free; doc stays a plain string; small transform core | Transform code is correctness-critical and version-locked; long offline degrades; useless for P2P/E2E |
| All of a doc's traffic on one session server | The serialization point that makes OT small; per-doc load is provably tiny | Server death interrupts its rooms (seconds); needs directory + epoch fencing; giant-audience docs need a viewer tier |
| Durable log append before ack/broadcast | "Seen implies saved" — survives any crash | Log write sits on the latency path; the log store must be fast and replicated |
| Snapshot every ~1000 ops + replay tail | Fast opens and recovery without rewriting the doc constantly | Tunable but real write amplification; a broken snapshotter silently inflates recovery time — must be monitored |
| Op log doubles as history; named versions = tags | A headline feature for zero marginal design | Log retained forever (cheap but real); restore must append, never rewrite |
| Per-op permission checks | Revocation takes effect mid-session | ACL cached in room memory + invalidation event — one more cache to keep honest |
A junior asks: "Why can't Google Docs just apply everyone's edits in the order they arrive at the server?" Explain the problem and the fix in five or six sentences.
Because an edit says "insert at position 4," and position 4 only means something in the exact document the editor was looking at when they typed. Your screen applies your own edits instantly — before the server sees them — so by the time someone else's edit reaches you, your document has shifted and their positions point at the wrong characters; applied blindly, every screen corrupts differently and they never agree again. The fix is transformation: before applying a concurrent edit, adjust its positions to account for the edits it hasn't seen — if you deleted four characters at the start, an incoming "insert at 4" becomes "insert at 0." To keep this manageable, all edits for a doc flow through one server that puts them in a single numbered order, so everyone adjusts against the same short list of recent edits. Every client's screen is just "last confirmed server state, plus my few unconfirmed edits," and every arriving edit gets reconciled through that little queue. That's why all screens end up with the identical text even though everyone typed at once.
Requirement flip: documents must be end-to-end encrypted — the server can never read content. Which parts of this design survive, which die, and what replaces them?
The killer: a server that can't read ops can't transform them, so server-side OT dies — and so do server-transformed comment anchors and any server-computed snapshot. What survives: the session server still assigns a total order to opaque encrypted blobs, still fences with epochs, still appends durably — ordering and storage don't require reading. Merge logic must move entirely to clients, and since clients may reconcile in different orders relative to their own pending edits, order-independent merging becomes the safe bet: this flip is precisely the stated trigger to switch to a CRDT, with per-character IDs living inside the encrypted payload. Snapshots become client-produced encrypted checkpoints. This is a great interview answer in miniature: the requirement change flips the crux decision, and you can say exactly why.
10x one room: a livestreamed lecture doc has 5 editors and 100,000 read-only viewers. The session server melts trying to broadcast. Sketch the viewer tier: what do relays subscribe to, what happens when a relay dies, and why don't viewers need OT?
Viewers hold no pending local edits, so their state is exactly "snapshot + ordered ops" — pure replay, no transformation, no two-queue model. Build a relay tree: the session server publishes each sequenced op once to a handful of tier-1 relays; each relay fans out to ~10K viewer sockets (100K viewers ≈ 10–15 relays, or two tiers at 1M). New viewers fetch snapshot + tail from storage, then subscribe at a relay from their current seq. A dying relay is a non-event: its viewers reconnect to a sibling and request the gap since their last seq — the same replay-from-seq recovery editors use. Editors stay pinned to the session server; the write path never learns the audience exists. One extra credit line: demote a viewer→editor transition to "reconnect to the session server as an editor," keeping the tiers clean.
Estimation drill: a busy doc takes 400 ops/sec in a worst-case burst. Snapshots cost a full 2 MB document write; replaying one op costs ~10 µs. Compare snapshot-every-100-ops vs every-10,000: steady-state write amplification and worst-case open/recovery replay time. Then commit to a policy.
Every 100 ops at 400 ops/sec = 4 snapshots/sec = 8 MB/sec of snapshot writes for one doc — pure amplification (the ops themselves are ~24 KB/sec). Worst-case replay: 100 ops × 10 µs = 1 ms. Every 10,000 ops: a snapshot every 25 sec (~80 KB/sec amortized — fine), but worst-case replay 10,000 × 10 µs = 100 ms — noticeable at open, painful when a server recovering 30K rooms multiplies it. The senior move is refusing the fixed constant: snapshot on ops-since-snapshot ≥ 1000 or doc idle for 30 sec, whichever first. Bursty docs bound their tail at ~10 ms replay; the idle rule means the common case (someone stopped typing) recovers from a snapshot that's seconds fresh, nearly free. Costs stay trivial against the numbers, which is the point of running them.
Now add "suggesting" mode (tracked changes): edits appear as proposals that a reviewer accepts or rejects later. What does a suggestion become in this machinery, and where's the subtle bug lurking?
A suggestion is an op tagged suggested_by that renders as markup instead of applying plainly — an insert shows as underlined proposed text (it does occupy positions), while a suggested deletion must NOT remove characters yet: it becomes an annotation range over the doomed text, transforming with ops exactly like a comment anchor. Accepting emits a new real op (the deletion, attributed to the accepter); rejecting emits an op removing the proposal markup. The log stays append-only and the review trail is itself history. The subtle bug: concurrent edits inside a pending suggested-deletion range — someone edits text that a pending suggestion wants to remove. Your anchor-transform rules must decide (grow the range? split it? invalidate the suggestion?), and whichever rule you pick must be identical on every client and the server, or screens will render different suggestions from the same log — divergence through the back door of the annotation layer.
What breaks if the directory grants doc 7 to server B while slow-but-alive server A still holds sockets and buffered ops — and a client is connected to each? Walk the exact save that prevents divergence.
Both servers accept ops and try to append at seq 2318 — without protection, whichever storage node each talks to might accept, and two clients now see different "truths": permanent divergence. The save is the epoch fence: B's grant came with epoch 8; A holds epoch 7; the log's append path atomically checks epoch ≥ highest-seen and A's append is rejected the moment B (epoch 8) writes anything. A's client's op is thus never durable and never acked — the client still shows it pending, reconnects (A drops sockets on seeing the rejection), lands on B, replays the gap from its last-acked seq, transforms its pending op over the missed ops, and resends with the same op_id. Every promise holds: nothing acked was lost, nothing applies twice, one order exists. The sentence that earns the point: the directory is allowed to be wrong; the log's fence is what's not allowed to be wrong.
"Design a distributed cache." You smile and reach for the obvious word, and the interviewer cuts you off: "No Redis. No Memcached. You're the team that builds the cache — every other team at this company is your customer. Show me what's inside the box."
This is the infrastructure flavor of the interview, and it's a favorite for senior loops at Meta and Google because it flips the usual game. Every chapter until now has let you draw a box labeled "cache" and move on. Now the box is the whole question. Everything you learned in chapter 1.6 (caching) and chapter 1.9 (consistent hashing) gets cashed in at build depth: not "use a hash ring" but which data structure holds the ring, who tells ten thousand clients it changed, and what happens to the SET that was in flight while it changed.
Here's the good news, and it's also the secret to the whole problem: a cache is the one distributed system that's allowed to forget. Losing data in a database is a career event. Losing data in a cache is a miss — the next read fetches it again. Every design choice in the next forty minutes gets easier if you say that out loud early and then actually design like you mean it.
Scope it out loud: "I'll build a shared, in-memory key-value cache for one datacenter. Operations: GET, SET, DEL, with per-key TTL — a time-to-live after which the key expires on its own. Values are opaque bytes up to 1 MB. Out of scope: rich data structures like sorted sets, cross-region replication, and persistence to disk. Fair?"
Then the non-functional list, with the honesty clause up front:
The backing database holds 30 TB. We obviously don't cache all of it — so how much? Access logs on almost any consumer workload show a steep skew, the 80/20 shape from chapter 2.6: a small slice of keys takes almost all the reads. Commit to a number: "Assume the hottest 10% of data serves roughly 90–95% of reads — I'd verify with a week of logs, but that's the planning number." So the working set is 3 TB of values.
Now the overhead nobody estimates: with 1 KB average values that's 3 billion entries, and each entry carries a key (~40 bytes), two list pointers, an expiry timestamp, a version — call it 100 bytes of metadata, +10%. Then fragmentation (we'll get honest about that later) means I only budget about half of each box's RAM for data. On 128 GB machines at 64 GB of entries each: 3.3 TB / 64 GB ≈ 52 — call it 56 primaries, and one replica each makes 112 nodes.
Check CPU: 2M GETs/sec over 56 primaries is ~36K ops/sec per node. A single-threaded event loop — one thread servicing all sockets, the design Redis proved — comfortably does 100K+ simple ops/sec, because a GET is a hash lookup and a pointer splice. Conclusion, said out loud: "Node count is set by RAM, not CPU. The cores will be bored."
Check latency, because it makes a decision: in-datacenter round trip is 100–500 µs; the lookup itself is ~1 µs. The network hop is the latency budget. Add a proxy tier in the middle and you've doubled p99 to serve an architecture diagram. So the hot path gets exactly one hop — that's a routing commitment, and estimation just made it for us.
One more number for the failure section: lose one primary with no replica and 1/56 ≈ 1.8% of the working set vanishes — which sounds tiny until you notice it turns into ~34K extra database reads/sec, a +34% database load spike, instantly. Write that down; it's why replicas exist here.
The wire protocol should be RESP-class: a simple length-prefixed text framing (RESP is Redis's protocol — "here come 3 fields, first is 3 bytes: SET…"). Trivial to parse, debuggable with netcat, and it makes pipelining natural: a client writes many commands into one socket without waiting for each reply, and replies come back in order. Pipelining is how one connection does tens of thousands of ops/sec despite a 300 µs round trip.
GET key [STALE_OK] -> VALUE v ver=n | MISS lease=t | MOVED shard=9 epoch=13
SET key v EX 60 [IFVER n | LEASE t] -> OK | REFUSED
DEL key -> OK
Three flags in that listing are doing quiet heavy lifting, and each is a promise I'll pay for in a deep dive: lease (stampede protection), IFVER (a versioned SET that refuses to overwrite newer data), and MOVED … epoch (how a client learns the cluster changed shape). The data model per entry: key → value bytes, expiry time, version, plus the two pointers that make eviction O(1).
Read path: the app calls get("user:42"). The client library hashes the key, binary-searches its local copy of the ring, opens (or reuses) a connection to the owning primary, and sends the GET. Hit: value back in ~300 µs. Miss: the app reads the database and SETs the result — cache-aside, the contract from chapter 1.6, now with teeth we'll add later.
Write path: the app writes the database first, then DEL or versioned-SET the key. The primary applies it and streams it asynchronously to its replica. No fsync, no quorum: it's a cache. If the primary dies mid-stream, the replica is a few milliseconds stale, and the worst outcome is a brief window of misses or slightly old values — which the TTL already bounds.
Control plane: a three-node config service (built on the consensus machinery from chapter 2.3 — this is the one place consensus belongs) stores the authoritative cluster map: the ring, each shard's primary and replica, and a monotonically increasing epoch number stamped on every version of the map. Clients cache the map and subscribe to changes. The config service is on zero data paths; if it dies, the cluster keeps serving with the last map — you just can't change topology until it's back.
Agenda, announced: "Three deep dives: the single-node engine, then the cluster under membership change — that's the hard one — then hot keys and stampedes."
A cache node is two data structures pointing at the same objects: a hash table for finding entries, and a doubly-linked list for ordering them by recency. Every GET does the dance: look up the entry in the map — O(1) — then unlink it from wherever it sits in the list (its two neighbors' pointers splice around it) and relink it at the head — O(1), because a doubly-linked entry knows both neighbors. Eviction pops from the tail: the least recently used entry, found in O(1), removed from list and map in O(1). No scans, ever. That's why the list must be doubly linked: with a singly-linked list, removing an arbitrary node means walking to find its predecessor — O(n), dead on arrival at 3 billion entries.
A coat check with a moving rail. The ticket book is the hash table: your number takes the attendant straight to your coat, no searching. The rail is the recency list: every time a coat is touched it slides to the front, so the dusty end of the rail is, by construction, the stuff nobody has wanted for the longest time. When the rack is full, the attendant doesn't deliberate — grab whatever's at the far end. The analogy breaks in one place: a real attendant walks the rail; our "rail" is pointers, so front and back are equally instant.
TTL: lazy plus active. Expiry the lazy way: every GET checks the entry's expiry timestamp; if it's past, treat it as a miss and delete. Free and correct — but a key that's never read again squats in memory until eviction finds it. So add an active sampler, and here I'll steal Redis's documented algorithm outright: ten times a second, pick 20 random keys from those carrying TTLs, delete the expired ones, and if more than 25% were expired, immediately sample again. That loop is self-tuning — it works harder exactly when dead keys pile up, and stays cheap when they don't.
Eviction: approximately LRU, on purpose. Here's a confession that reads senior: perfect LRU at this scale is a lie you stop telling. A single global recency list means every GET on every connection mutates one shared structure — on a multi-threaded node it's a lock brawl, and the list pointers alone cost 48 GB across our 3 billion entries. Redis's documented answer: don't keep a perfect order. Store a last-access clock per entry, and on eviction sample a handful of random entries (default 5) and evict the oldest of them, keeping a small pool of best candidates between rounds. The measured behavior is nearly indistinguishable from true LRU on real workloads — because you don't need the coldest key, just a cold one. Policy knob: allkeys (anything may be evicted — right for a pure cache) versus volatile (only TTL-carrying keys are evictable — for mixed workloads where some keys must not vanish). I'll run allkeys-LRU.
Memory honesty. Ask any allocator for 3 billion variably-sized objects with constant churn and you get fragmentation: the heap's RSS (what the OS says you use) creeps far above the bytes you think you stored. Slab allocation is the classic defense — carve memory into fixed-size classes (64 B, 128 B, 256 B…) and round each value up to its class, trading internal waste for immunity to scatter. I'll use a slab/arena design and monitor the fragmentation ratio, RSS over logical bytes, and treat 1.5 as an alarm. But slabs have their own famous failure mode:
Twitter ran memcached forks on thousands of instances, and in 2012 open-sourced their own build, Twemcache, with an engineering post explaining why. The villain was slab calcification. Memcached hands out memory in 1 MB pages, each page committed to one size class — and in that era, once a page joined a class, it stayed there. Twitter ships a product change; serialized objects get bigger; the new size class owns almost no pages, so those writes evict each other furiously while gigabytes sit locked in the old, now-idle classes. Hit rate craters on a box that is technically full of free-ish memory. Twemcache's fix was blunt and effective: random slab eviction — evict an entire randomly chosen 1 MB page, wiping the items on it, and reassign the page to whatever class needs it now, keeping memory fluid as the workload drifts. Twitter's cache team later went further and built Pelikan, a from-scratch modular cache server, to escape their pile of forks. The lesson survives in any build-a-cache answer: your allocator's shape is workload-dependent, workloads drift with every deploy, and a cache that can't rebalance memory between size classes ages like concrete.
This is the room that's on fire. A static cluster is easy — hash, modulo, done. The question exists because clusters change: nodes die, capacity gets added, deploys roll. Naive hash(key) mod N means going from 56 to 57 nodes remaps ~98% of keys — hit rate falls off a cliff, the database eats 2M reads/sec, and your cache just caused the outage it existed to prevent. And even with a good ring, there's a correctness trap inside the transition itself: two nodes both believing they own a key, a zombie primary accepting writes after failover, a client routing on a map that's ten seconds stale. Cluster membership change — who owns what, who knows who owns what, and what happens to requests in flight while the answer changes — is where this problem is won.
The ring, concretely. Chapter 1.9 gave you the picture; now build it. Hash every node into the same 64-bit space as the keys. The ring is just a sorted array of (hash, node) pairs; routing a key is hashing it and binary-searching for the first node hash clockwise — O(log V) over a few thousand entries, nanoseconds, done inside the client library. Adding a node inserts its positions into the array and steals only the arc segments just counter-clockwise of them: K/N keys move — 56→57 nodes touches ~1.8% of the working set, a hit-rate ripple from 95% to maybe 93% for a few minutes, which the database's headroom absorbs.
Virtual nodes fix the two problems a naive ring keeps: with one position per node, random placement makes some arcs 3× the size of others, and a dead node dumps its entire arc on one neighbor. So each physical node hashes to ~200 positions. Law of large numbers: 200 samples per node tightens load spread to a few percent — and a dead node's 200 slivers scatter across all survivors, each absorbing a crumb instead of one neighbor swallowing a boulder.
Routing: I'll commit to smart clients. Three candidates. A proxy tier (the mcrouter/Twemproxy shape) centralizes routing logic behind a dumb client — but it adds the network hop our 1 ms p99 can't afford, plus a fleet to operate. Gossip mode (Redis Cluster's shape) lets any node redirect you to the right one — no extra component, but now every cache node runs membership protocol, and debugging "who thought what, when" during an incident is genuinely painful. A smart client — a library embedding the ring, subscribed to the config service — gives one-hop latency and a dumb, boring server. The cost, named honestly: the client library is now versioned software running inside every app, in every language your company uses, and a routing bug ships at the speed of your slowest team's deploys. At our scale (one datacenter, a couple of languages) that's the right trade; at Facebook's polyglot scale, the proxy wins — which is exactly why mcrouter exists.
Topology change, the protocol. Every published map carries an epoch. Every client request carries the epoch it routed with; every node knows the epoch of the map that made it an owner. Node D joins: the config service publishes epoch 13; D's designated ranges start warming — D pulls those keys from their old owners with a cursor scan while old owners keep serving; when a range is warm, epoch 14 flips ownership. A client still on epoch 13 sends GET to the old owner and gets back MOVED shard=9 epoch=14 — a redirect that doubles as a nudge: refresh your map. The client re-fetches, retries against D, one extra round trip, no error surfaced to the app. Requests never fail during rebalance; they occasionally take two hops. And if warming fails halfway? Abort by publishing the old ownership at epoch 15 — the keys D half-copied are just duplicates that will TTL out. Forgetting is free. Say that sentence in the interview; it's the difference between this problem and building a database.
Replication and failover, with the fence. Each primary streams its writes to one replica, async, no acknowledgment on the hot path. The replica's job is not durability — it's hit-rate insurance: when the primary dies, the config service health-checker notices (three missed heartbeats, ~5 seconds), promotes the replica, bumps the epoch, publishes. Clients fail over in one map refresh. The nasty case is the primary that isn't dead — a 20-second GC pause, a network partition — and comes back believing it still owns shard 7. That's split-brain, and the guard is the fencing pattern from chapter 2.3: every write between peers and from clients carries the epoch; a node that sees traffic stamped with a higher epoch than its own demotes itself; peers and clients refuse anything stamped lower. The zombie can talk; nobody with a current map will listen.
And the honest accounting, volunteered before the interviewer asks: async replication means the promoted replica is missing the last few milliseconds of writes. In a database that's data loss and a scandal. Here it's a handful of keys that read slightly stale or miss, then self-heal from the source of truth. Designing a cache means noticing, over and over, that the forgiving failure model is the whole reason the design can stay this simple — sync replication would double write latency to protect data that never needed protecting.
Sharding spreads keys, not load — chapter 2.6's lesson, now on our own hardware. When a celebrity posts, feed:cr7 gets 500K reads/sec, and every one of them hashes to the same primary that was sized for 36K. The ladder, in order of reach:
STALE_OK GET may go to primary or replica at random — doubling read capacity for keys where milliseconds of replication lag don't matter.feed:cr7#1…#8) that land on different shards; readers pick a random suffix. Writes cost 8×; reads scale 8×.Now the stampede — the trap from chapter 1.6, which we can finally fix inside the cache instead of asking every application team to fix it politely in their code. A popular key expires. In the next 50 milliseconds, a thousand app servers miss simultaneously, and all thousand send the same expensive query to the database. The cache caused a denial-of-service attack on its own backing store.
The office fridge runs out of milk. Without coordination, all twelve people who notice independently drive to the store, and twelve cartons of milk arrive. The fix is a sticky note on the fridge: "gone for milk — back in 10." The first person to notice writes it; everyone after reads the note and either waits or drinks their coffee black for now. The note is a lease, black coffee is the stale value, and the store is your database. The analogy's one gap: our fridge hands out the note itself, atomically, so two people can never both believe they left first.
Build the note into the protocol. On a miss, the cache doesn't just say MISS — it says MISS lease=7291 to the first caller, and for the next few seconds tells everyone else WAIT (or hands them the just-expired value if the GET carried STALE_OK — stale-while-revalidate as an explicit API flag, so the caller chooses whether yesterday's milk is acceptable). One client recomputes; its SET … LEASE 7291 is accepted; a thousand waiting clients get hits. Database sees one query instead of a thousand.
STALE_OK get the just-expired value instead of waiting. The token also acts as a version guard: a DEL invalidates outstanding leases, so a late SET from a stale reader is refused.That last clause deserves its own minute, because it quietly fixes the most famous correctness bug in caching. Walk the race: reader R misses on user:42 and reads value v1 from the database. While R is stuck in a GC pause, writer W updates the row to v2 and issues the cache-aside DEL. Then R wakes up and SETs v1. The cache now serves stale v1 until TTL — the write happened before the delete, so the delete cleaned nothing. With leases: R's SET carries lease 7291, but W's DEL invalidated that lease, so the cache answers REFUSED and R's stale write dies at the door. Same protection via IFVER for systems that carry a real version (a row version, a commit timestamp): the cache keeps the highest version and refuses regressions — last-writer-wins decided by the data's clock, not the network's race. For belt-and-braces at company scale, add a CDC invalidation channel: tail the database's change log (chapter 2.9) and publish DELs/versioned SETs from the log itself, so invalidation doesn't depend on every application team remembering to send the DEL.
Almost everything in this deep dive is documented in one famous paper: "Scaling Memcache at Facebook" (NSDI 2013), describing the fleet that fronted Facebook's MySQL tier — billions of requests per second against trillions of items. Leases are their invention, built to kill two birds: thundering herds and stale sets. On a miss, the memcached server hands the client a 64-bit lease token bound to the key; only a SET carrying that token is accepted; a delete arriving in between invalidates the token, so the stale-set race dies exactly as described above. The server rate-limits token issuance to once per ten seconds per key, telling later clients to wait or use slightly stale data. In one measurement in the paper, turning leases on dropped peak database query rate from about 17,000 to 1,300 per second. Their second trick is the gutter: a small idle pool, about 1% of the fleet, that takes over key traffic when a memcached server dies — with short TTLs, so gutter entries expire fast instead of needing invalidations. The alternative — rehashing a dead server's keys onto its neighbors — risks toppling them in sequence when hot keys land badly; the gutter absorbs the blast in a padded room instead. The client-side routing logic in this architecture became mcrouter, open-sourced in 2014 — the proxy road we consciously didn't take at our smaller scale.
Node death is the rehearsed scenario: replica promoted in ~5 seconds via epoch bump; hit rate dips a point or two while the new primary warms its own replacement replica. The rule that makes this survivable is the N−1 rule: the database must be provisioned to survive the cache's worst single failure — the largest shard dark, cold, and stampeding. If your database only survives when the cache is at full strength, the cache isn't protecting the database; it's holding it hostage. For a whole-rack event, borrow Facebook's gutter: a couple of idle nodes that take orphaned traffic with 10-second TTLs.
Rollouts: a cache restart discards its memory, so a naive rolling deploy is a self-inflicted 56-shard cold-start. The deploy dance: bring the new-version process up as an extra replica, let it warm from the stream, promote it (epoch bump), retire the old primary. Hit rate never dips; the fleet upgrades one epoch at a time. Watch per-shard hit rate between steps and halt on regression.
Dashboards, in priority order: per-shard hit rate (the product metric), database read QPS (the thing hit rate exists to protect — page on this), evictions/sec and expired-vs-evicted ratio (evictions rising with flat traffic means the working set outgrew RAM), fragmentation ratio (RSS/logical — Twitter's calcification shows up here first), replication lag, hot-key top-N, p99 per command per shard, and config-epoch age per client (a client stuck three epochs behind is a routing bug waiting to page you).
At 10x: working set 33 TB → ~560 primaries. The ring and config service barely notice — topology metadata is tiny. Two things break first: every app server now holds hundreds of connections (the smart-client trade-off coming due — I'd revisit a sidecar proxy per host, which is mcrouter's actual job), and RAM economics start arguing for a tiered cache — hot tail in RAM, warm tail on NVMe — which changes the single-node engine but, satisfyingly, not the cluster protocol. The epoch machinery we built is scale-independent; that's the sign it was the right crux.
How interviewers probe this problem once your diagram is up:
| Decision | Chosen because | What it costs |
|---|---|---|
| Smart client, one hop | p99 budget is one network RTT; server stays dumb | Routing logic versioned into every app, per language; connection fan-out grows with cluster size |
| Consistent hashing + ~200 vnodes | Node change moves ~K/N keys; dead node scatters into crumbs | Ring metadata and warm-up choreography; slightly uneven load vs a perfect slot table |
| Async replication, no fsync | It's a cache — loss = misses; hot path stays sub-ms | Milliseconds of writes lost on failover; brief bounded staleness |
| Epoch-fenced config service | One authority for topology; zombies fenced by number, not by hope | A consensus service to operate; topology changes gated on its availability |
| Sampled LRU, allkeys | No global lock, tiny metadata, ~true-LRU hit rate | Occasionally evicts a warm key an exact LRU would have kept |
| Leases + stale-while-revalidate in the protocol | Stampedes and stale-set races die in one mechanism, for every client team at once | Protocol complexity; brief windows where clients see WAIT or stale values |
Explain to a junior, in five or six sentences, why we route with a hash ring instead of hash(key) mod N — and why each node appears on the ring 200 times.
With mod N, the divisor is baked into every key's address, so changing the node count from 56 to 57 re-addresses almost every key at once — the cache goes effectively cold and the database takes the full read load. A hash ring fixes the addresses: keys and nodes hash into the same circle, each key belongs to the next node clockwise, so adding a node only claims the thin arcs just behind its positions — about 1/57th of the keys move and everything else stays put. But one position per node makes luck matter: random placement gives some nodes triple-sized arcs, and when a node dies its whole arc lands on a single neighbor. So each node appears at 200 points: averaging 200 slivers evens the load to within a few percent, and a dead node's slivers scatter across all survivors instead of crushing one. Small keys moved on change, even load, gentle failure — that's the whole trick.
The working set grows 10x to 33 TB but traffic only 3x. Re-run the sizing: how many primaries, what breaks first in the current design, and what would you change before it does?
33 TB × 1.1 overhead ≈ 36 TB / 64 GB per node ≈ 560 primaries, ~1,120 nodes with replicas — RAM still sets the count, since 6M GETs/sec over 560 nodes is only ~11K ops/sec each. First break: the smart client's connection fan-out — every app server holding ~1,120 connections, times thousands of app servers, is millions of sockets; fix with a per-host sidecar proxy that multiplexes (the mcrouter shape — the trade-off we consciously deferred has come due). Second: RAM cost at 36 TB argues for tiering — RAM for the hot tail, NVMe for the warm tail — which rewrites the single-node engine but leaves the ring, epochs, and lease protocol untouched.
A 100-node cluster grows to 110. What fraction of keys move (a) with consistent hashing, (b) with hash mod N? And why do virtual nodes matter more for node removal than addition?
(a) Each new node claims about 1/(current size) as it joins; net effect ≈ 10/110 ≈ 9% of keys move — hit rate ripples a few points and recovers. (b) Mod N: a key stays put only if hash mod 100 = hash mod 110, which holds for roughly 1 in 11 keys — about 90% move, a synthetic cold start. Removal is where vnodes shine: an added node politely pulls small arcs from many donors, but a removed node's load lands instantly wherever the ring says. With one position, that's one neighbor absorbing an entire node's keys and traffic — likely a cascading overload. With 200 positions, the dead node's ranges scatter across all survivors at ~1% extra each.
After Tuesday's deploy, cluster hit rate fell from 95% to 81%. Traffic is flat, no nodes died, TTLs unchanged. Evictions/sec tripled. Walk your diagnosis.
Evictions up with flat traffic and healthy nodes means entries are fighting for memory they used to have — something about the data changed. Check fragmentation ratio and per-size-class stats: the classic culprit is the deploy changing serialized value sizes (new field, new format), landing writes in slab classes that own almost no pages — Twitter's slab calcification. New-size entries evict each other furiously while old classes sit on idle memory. Confirm: eviction age (evicted entries that are seconds old = memory starvation, not natural aging) and value-size histogram before/after deploy. Fix: slab rebalancing (evict and reassign whole pages, Twemcache-style) or roll the fleet to reset allocator state — via replica-promote, never a cold restart. Senior tell: you asked "what changed?" before touching the cluster.
Requirement flip: product now demands that after a user edits their settings, no server ever reads a settings value older than 1 second from the cache. Your TTLs are 10 minutes. What do you change, and what does it cost?
First say the honest thing: plain cache-aside can't promise this — a DEL can be lost, a race can re-insert stale data, TTL is the only backstop. To get a real bound: (1) versioned SETs so stale writes are refused (kills the race), plus (2) a CDC invalidation channel tailing the database log, so invalidation is guaranteed by the log rather than by app-team discipline — end-to-end lag budget under 1s, monitored as its own SLO. Also disable the client L1 for the settings keyspace (its 1s staleness now eats the entire budget). Costs: the CDC pipeline is a new stateful dependency whose lag can now violate a product promise; and you should ask whether this keyspace belongs in the cache at all — settings reads at settings-write frequency might just go to the primary DB. Scoping a guarantee down to one keyspace instead of re-architecting the whole cache is the senior move.
One key — a live-match scoreboard — hits 2M reads/sec, your entire cluster's read budget, while being written twice a second. Walk the mitigation ladder with numbers.
The owning primary rated for ~36K ops/sec gets 2M — 55x over. Ladder: (1) client L1 with a 500ms TTL: 2,000 app servers each refresh at most twice a second → ≤4K reads/sec reach the cluster, a 500x cut, staleness bounded at 500ms — acceptable for a scoreboard and probably sufficient alone. (2) If not: key fan-out ×8 (score:match#1…#8) spreads residual reads across shards; writes go to all 8, trivial at 2/sec — fan-out suits exactly this read-huge/write-tiny shape. (3) Replica-random reads with STALE_OK for another 2x. Note what you did NOT do: reshard the cluster or buy bigger boxes — a single hot key is a routing problem, not a capacity problem; the fix belongs at the client edge, not the fleet.
Forty minutes into a Meta loop, you've drawn a queue into three different designs. Notifications went through one. Click aggregation went through one. The outbox in the payment design went through one. Each time you said "and Kafka absorbs the spike here," and the interviewer nodded.
Then they stop nodding. "You've leaned on that box four times today. Let's build it. Design Kafka." The comfortable primitive becomes the whole problem — because now you are the thing that must not lose the message.
It looks like "design a queue" and it isn't. A queue delivers a message and forgets it. Kafka is a replicated append-only log — data stays put after it's read, many independent readers walk it at their own pace, and any of them can rewind to last Tuesday. Chapter 1.11 taught you to use a log; this chapter builds one.
Scope out loud: "I'll build the broker cluster — topics, partitions, replication, producers, consumer groups, retention and compaction. Out of scope: stream processing, schema registry, connectors, cross-cluster mirroring. Fair?"
Functional: publish to a topic, split into partitions; records inside one partition are strictly ordered and carry a monotonically increasing offset (a position, not an invented ID); consumers read from any offset, so history is replayable; consumer groups split a topic's partitions among cooperating processes; retention by time or size, plus compaction for keyed topics; producer-selectable durability via acks=0/1/all.
Non-functional, with numbers I'm committing to — these decide the machine:
acks=all survives any one broker dying.acks=all.Note the requirement I did not write: global ordering across a topic. Refusing it is correct engineering, and I'll show why.
Storage first, because it's the surprise. 1 GB/sec × 86,400 = 86 TB/day. Seven days is ~600 TB; at RF=3 that's 1.8 PB on disk. Give each broker twelve 2 TB drives as JBOD — 24 TB raw, and never plan past ~60% full or compaction and recovery have nowhere to breathe — so ~14 TB usable. 1,800 ÷ 14 ≈ 125 brokers. Call it 130.
Now network, and watch it not matter. Each broker leads ~8 MB/sec of new writes, ships them to 2 followers (16 MB/sec out), takes ~16 MB/sec as a follower for others, and serves 3 GB/sec ÷ 130 ≈ 23 MB/sec to consumers. Call it 24 MB/sec in, 39 MB/sec out — about 3% of a 10 Gbps NIC.
Say the conclusion out loud, because it's the architecture: "This cluster is disk-capacity-bound, not throughput-bound. The biggest cost lever in the design is retention — dropping the default from 7 days to 3 takes me from 130 brokers to about 55. So retention becomes a per-topic decision with an owner, not a config nobody looks at."
A partition is a spiral notebook on a shop counter. The shopkeeper only ever writes on the next blank line — never in the middle, never erasing. Every customer keeps a sticky note marking the line they've read up to, and moves it at their own speed; one slow customer slows nobody, and a new customer can start from page one. One writer appending, many readers with independent bookmarks. Where it breaks: real notebooks don't bin their oldest pages after seven days. Ours does, and that retention rule is what sizes the cluster.
Most candidates reach for "keep hot data in memory" as if disk were the enemy. Kafka's founding insight is that a log has the one access pattern disks are good at, and that the OS already runs the cache you were about to rebuild. Two facts do the work. Appends are sequential. A disk array that manages a pitiful few hundred KB/sec of random writes will do hundreds of MB/sec of linear writes — a gap of thousands of times, and the same shape holds on SSDs. A log never seeks. The page cache is already the cache. A batch the broker writes lands in page cache; a consumer a few seconds behind finds it still in RAM.
Put our numbers on it. A broker has 64 GB of RAM. We give the JVM a deliberately small heap — 6 GB — because the broker caches nothing itself, leaving ~55 GB of page cache. New bytes land at ~24 MB/sec. 55 GB ÷ 24 MB/sec ≈ 2,300 seconds, about 38 minutes of log resident in RAM. Two decisions fall out:
One more free win: because the broker never parses the bytes it stores, it hands a byte range from page cache straight to the socket with sendfile — zero-copy. The normal path copies four times (disk → page cache → app buffer → socket buffer → NIC); sendfile removes two copies and a syscall, and the data never enters the JVM heap. Volunteer the caveat: TLS kills zero-copy, because encryption happens in userspace, and encrypted clusters cost real CPU.
Produce(topic, partition?, key?, value, acks) → {partition, offset}. No partition means the client hashes the key; no key means round-robin.Fetch(topic, partition, offset, max_bytes) → a run of record batches. Followers use this exact call — replication is a consumer that happens to be a broker.Metadata(topics) → leader, replica set and ISR per partition. Clients cache it, refresh on a "not the leader" error.CommitOffset(group, …) / FetchOffset(group, …); admin CreateTopic(name, partitions, rf, retention_ms, cleanup_policy).There's no database — the data model is the file layout. One directory per partition (orders-7/) holding segment files named for the offset they start at: 00000000000000368000.log, plus .index (offset → byte position) and .timeindex. A batch carries a base offset, record count, CRC, compression codec, producer ID and base sequence number, then delta-encoded records. Two details matter later: the batch is the unit of compression and transfer, and producer ID plus sequence is what makes retries safe.
Write path. The producer fetches metadata, hashes the key to pick a partition, and appends to an in-memory batch. It waits up to linger.ms to fill the batch, compresses it once, and sends it to that partition's leader only. The leader appends to the tail of its active segment — a page-cache write, no fsync per message; durability comes from replication, not from forcing bytes to a platter. Followers, permanently fetching, pull the new bytes; once every in-sync replica has them, the leader advances the high watermark and acks the producer.
Read path. A consumer sends Fetch(partition, offset) to the leader. The broker binary-searches segment filenames, uses the sparse index to land near the offset, scans forward to the exact record, and sendfiles a range of the file to the socket. The consumer processes, then commits its offset — to a Kafka topic, because everything here is a log. Nothing is deleted on delivery: that's the difference between a log and a queue, and it's why replay is free.
Agenda, announced rather than waited for: "Three things deserve depth — the storage engine, replication and the guarantees hanging off it, and consumer groups. Replication is the real difficulty, so it gets most of our time."
Each partition is a directory, and its log is cut into segments of about 1 GB; exactly one is open for appending, the rest immutable. Segments exist because you need a unit of deletion — you can't truncate the front of a file, but deleting a whole 1 GB file is instant.
The index is where candidates get vague, so be concrete. Consumers ask for offsets, and an offset is not a byte position. A dense offset→byte index would be enormous and would defeat the cheap append, so the index is sparse: one entry per ~4 KB of log, small enough to memory-map. The cost is a short forward scan — sequential reading from page cache, the thing we're fastest at.
Two kinds of log, and when each is right. Retention (delete) drops whole segments past 7 days or a size cap — right for event streams, where each record is a fact about a moment and a month-old fact carries no ongoing meaning. Compaction is for keyed state: if the topic means "the current shipping address for every user," you can never drop the last record for a key, or a consumer rebuilding from offset 0 gets a user with no address. Compaction keeps at least the latest value per key and garbage-collects the rest; a null value is a tombstone ("this key is deleted"), retained for a grace window so slow consumers see it.
The rule: compact when the topic is a table pretending to be a stream; delete when it's genuinely a stream of events. Kafka's own __consumer_offsets is compacted — latest offset per (group, topic, partition) is exactly a keyed table stored as a log. Two honest costs: compaction does real I/O rewriting segments, and it destroys strict replay, since you see only where each key ended up.
A partition does two jobs that pull against each other. It is the unit of ordering — strict within a partition, none across them — and the unit of parallelism: one leader, and within a consumer group exactly one reader.
The heuristic I'd commit to: partitions = max(target throughput ÷ 30 MB/sec, peak consumer instances), times ~2 for headroom. For our 1 GB/sec topic that's max(34, consumer count), and consumer parallelism almost always wins. Headroom matters because growing partitions later breaks key stickiness: hash(key) mod N changes with N, so a user's records start landing elsewhere and their old and new events lose relative order forever. Adding partitions to a keyed topic is a migration (chapter 2.11), not a config change — say that unprompted, it's the trap here.
Kafka came out of LinkedIn's data-integration mess around 2010: dozens of bespoke point-to-point pipelines shuttling activity data between databases, search, monitoring and Hadoop, each built and broken separately. Jay Kreps, Neha Narkhede and Jun Rao's 2011 NetDB paper, "Kafka: a Distributed Messaging System for Log Processing," is blunt about the unconventional choices: messages addressed by offset rather than an ID, which "avoids the overhead of maintaining auxiliary, seek-intensive random-access index structures"; no in-process message cache at all, relying on the OS page cache to avoid double buffering and keep garbage collection cheap; sendfile; pull-based consumers. At the time LinkedIn's Kafka carried "hundreds of gigabytes of data and close to a billion messages per day." In October 2019 they published the numbers again: 7 trillion messages per day across 100+ clusters, 4,000+ brokers, 100,000+ topics and 7 million partitions. Keep that last pair — roughly 1,750 partitions per broker in one of the largest deployments on earth. That's your empirical ceiling, not a made-up rule.
Everything else in Kafka is a well-chosen file format. The hard part is keeping N copies of an append-only log consistent while machines die, networks split, and a producer waits on an acknowledgement that has to mean something. Every guarantee reduces to three mechanisms: which replicas count as in-sync, what an ack waits for, and how a stale leader is stopped from writing after it's replaced. Get one wrong and you have a system that silently loses acknowledged data — worse than one that loses it loudly. This is where your interview time goes.
Each partition has one leader and R−1 followers. All client traffic goes through the leader; followers do nothing but continuously Fetch, using the consumer API. The consequence matters more than the elegance: a follower's position is just an offset, so the leader knows exactly how far behind each one is. From that comes the ISR — the in-sync replica set: replicas (leader included) that have fetched within replica.lag.time.max.ms, default 30 seconds. A follower that stalls on GC, a slow disk or a network hiccup drops out, and rejoins when it catches up. The ISR is dynamic, and that word is the whole trap.
The leader also tracks the high watermark: the highest offset every ISR member holds. Consumers may only read up to it — so unreplicated data on the leader is invisible, and no consumer can read a record a later failover could erase. Name that; it's a quietly excellent design choice.
acks=0: the producer writes to its socket and moves on. No promise — a dropped connection loses the batch and nobody finds out. Legitimate for exactly one class of data: telemetry you'd rather lose than slow the app for.
acks=1: the leader acks as soon as the batch is in its own log. Walk it concretely, because this is the scenario. A producer sends offsets 5,000–5,100. The leader appends and acks all of them; your app has moved on, dropped its copy, told the user "order placed." Followers are 40 ms behind, at 5,060. The leader loses power. The controller promotes a follower whose log ends at 5,060. Offsets 5,061–5,100 — forty acknowledged records — never existed as far as the cluster is concerned. No error, no retry, no alarm. That is the loss window: whatever the leader took but hadn't shipped.
acks=all: the leader waits until every ISR member has the batch. Failover then picks a replica that by definition holds your data. The cost is one extra round trip to the slowest in-sync follower — which is why we budgeted 50 ms, not 5.
You're dictating a phone number to a room of three people. acks=0 is saying it and walking out. acks=1 is waiting for one "got it" — and if that person's memory is the only copy when they leave, the number is gone. acks=all is waiting until everyone who's paying attention has written it down. Here's the break, and it's the bug: "everyone paying attention" is a shrinking list. If two of the three drift off and nobody tells you, "everyone wrote it down" quietly means one person did — the weak promise while you believe you have the strong one. The fix is a rule stated in advance: I refuse to dictate unless at least two people are listening.
min.insync.replicas and the unclean-election choiceThat's why acks=all alone isn't enough. With RF=3, if two followers fall behind, the ISR is just the leader, and acks=all now means "the leader has it" — acks=1 in a costume.
min.insync.replicas=2 closes it: below 2 in-sync replicas the leader refuses the write with NOT_ENOUGH_REPLICAS, so the producer gets an error instead of a false promise. RF=3 + acks=all + min.insync.replicas=2 is the standard durable configuration: one broker loss is invisible, a second turns into rejected writes rather than silent loss. That's consistency chosen over availability (chapter 1.10), per topic.
The other half is unclean leader election. All ISR members are dead; only an out-of-sync replica survives, missing thousands of records. Two options, no third. On: promote it — the partition returns immediately, the missing records are gone, and the log diverges, so offset 5,090 now means a different record than it did an hour ago. Off: the partition stays offline until an ISR member returns — producers error, consumers stall, someone gets paged, and not one acknowledged record is lost.
My commitment, per topic class: off for anything a business decision depends on — orders, payments, the chapter 2.1 outbox, audit trails; availability there is a page, data loss is a lawsuit. On for telemetry and clickstream, where a gap in a graph beats an offline partition blocking every producer in the fleet. Not a preference, a classification. Modern Kafka defaults to off, which is right — and it wasn't always.
When Kafka 0.8 introduced replication, Kyle Kingsbury tested it in his "Call me maybe" series and found exactly the failure the machinery above exists to prevent. His analysis puts it plainly: Kafka "attains this goal by allowing the ISR to shrink to just one node: the leader itself. In this state, the leader is acknowledging writes which have only been persisted locally." Partition that leader away from ZooKeeper and a replica arbitrarily far behind gets promoted. In one run: 987 writes acknowledged, 520 of them lost — more than half the data the system had already called safe, gone without an error. The response is what we just committed to: from 0.8.2 you could set min.insync.replicas so the leader rejects writes rather than acking alone, and unclean.leader.election.enable=false so a lagging replica is never promoted; later releases flipped that default to off. Two lessons for any interview: a durability setting is only as strong as the weakest state its quorum can reach, and "acknowledged" is a promise you must be able to trace to durable copies — otherwise it's a lie with good intentions.
One failure min.insync.replicas doesn't touch. Broker A leads P0; A's network drops; the controller elects B. Thirty seconds later A returns and, having been out of contact, still believes it's the leader. If A accepts a write now, or a follower fetches from it, you get two divergent versions of the same offsets. That's chapter 2.3's split brain, with chapter 2.3's answer: fencing tokens, called leader epochs here.
Every leadership change increments the partition's epoch, handed out only by the controller, and requests carry it. A stamped with epoch 5 against a cluster on epoch 6 is refused: it cannot append, and followers won't fetch from it. A then learns it's a follower, asks the real leader "where did epoch 5 end?", truncates to that point, discards what it accepted while isolated, and refetches. Without epochs, replicas could only compare high watermarks — not enough, after a rapid double failover, to find where the logs actually diverged. Silent divergence is the worst bug class in the system.
Something must hold the cluster's truth: which brokers are alive, who leads each partition, who's in each ISR. That's the controller, and there must be exactly one — chapter 2.3's leader election. Kafka rented that job to ZooKeeper for most of its life: it worked, but it meant a second distributed system with a second failure model, and on failover the new controller had to read full state out of ZooKeeper, which at hundreds of thousands of partitions took minutes. KRaft replaced it — 3 or 5 brokers run Raft among themselves and cluster metadata becomes, of course, a replicated log they all follow, so failover is a standby that already has the log taking over in seconds. KRaft went production-ready in Kafka 3.3 and ZooKeeper was removed in 4.0. Honest framing: not a clever new algorithm, just Kafka finally eating its own dog food.
Producers. Two settings carry the throughput. linger.ms makes the producer wait a few milliseconds so one request carries hundreds of records, and compression runs on the whole batch (lz4 or zstd), which beats per-record compression because records in a topic share structure. The batch stays compressed on disk and is served compressed — that's what keeps zero-copy alive. Trading 5 ms of latency for a 5–10x throughput gain is usually right; knowing you're trading is the point.
Now the correctness hole. A producer sends a batch, the leader writes it, the ack is lost on the way back, the producer retries, the leader writes it again. One record, two offsets — a duplicate created by the network. Chapter 2.2's problem, solved by moving chapter 2.2's answer into the protocol: the idempotent producer. Each producer gets a producer ID; every batch carries it plus a per-partition sequence number. The leader remembers the last sequence accepted per producer per partition, and a retry with a sequence it has already seen returns success without appending. It also preserves ordering under retries — without it, a failed batch retried behind a successful one reorders your log. On by default, effectively free.
Transactions, in one paragraph, because they'll ask. A producer registers a transactional ID with a coordinator, writes to many partitions, and commits; the coordinator writes commit markers, and read_committed consumers skip uncommitted or aborted records. Commit consumer offsets in the same transaction and you get real exactly-once for consume-transform-produce pipelines. What it is not: atomicity between Kafka and your database — that's still chapter 2.1's outbox. Costs: a coordinator in the path, lower throughput, and a stuck transaction blocking read_committed consumers behind it.
Why consumers pull is a chapter 2.7 payoff. A pushing broker would have to model every consumer's capacity and buffer for the slow ones — flow control rebuilt in the broker for every client. With pull, a consumer asks for what it can handle and stops asking when busy: backpressure is automatic, "falling behind" becomes a number you can graph, replay is trivial because the fetch carries the offset, and batching is natural. Honest cost: a caught-up consumer would spin, so brokers hold fetch requests open for up to fetch.max.wait.ms rather than answering "nothing."
Consumer groups. Members share a group.id and split the topic's partitions, each partition going to exactly one member. A group coordinator tracks membership by heartbeat, and a change triggers a rebalance.
Rebalances hurt more than people expect. The classic protocol is eager: everyone revokes everything, the leader recomputes, everyone resumes — the group stops consuming throughout, and per-partition in-memory state is thrown away. Worse, they cascade: a consumer quiet during a long GC pause gets evicted, and its rebalance's stop-the-world pause pushes another member past its heartbeat deadline. That's a rebalance storm. Cooperative (incremental) rebalancing only revokes the partitions that must actually move. Worth saying out loud that it is not automatic: the classic consumer still ships with partition.assignment.strategy set to range-then-cooperative-sticky, and because the range assigner is first in that list it wins — you get eager behaviour unless you remove it. Kafka 4.0's new consumer group protocol (KIP-848) makes incremental assignment the mechanism and moves the computation from a client member to the broker-side coordinator, but consumers opt in with group.protocol=consumer. Two things I'd do unprompted: set max.poll.interval.ms honestly for the slowest record I might process, and use static group membership so a rolling restart doesn't reshuffle for every pod.
Offset commits — the duplicate window. Consumers commit progress to __consumer_offsets, and the order of "process" and "commit" is the delivery guarantee. Walk it: the consumer polls offsets 100–199, processes 100–149 — each one charging a card — and the pod is evicted before committing. On restart the committed offset is still 100, so 100–199 are redelivered and 50 cards get charged twice. That's at-least-once. Flip the order — commit 199, then process — and a crash mid-batch means 150–199 are never processed and never redelivered: at-most-once, silent loss instead of duplicates.
I'll commit to at-least-once and make consumers idempotent, because a duplicate you can absorb beats a message you can't recover. Chapter 2.2 exactly: a natural key on the record plus a conditional write downstream turns "delivered twice" into a no-op. The window is bounded and knowable — everything between the last commit and the crash — so committing more often shrinks it at the cost of more writes.
Ordering, honestly. Kafka orders records within a partition; there is no ordering across partitions, ever. A globally ordered topic means one partition — one leader, one broker's disk, one consumer per group, a ceiling in the tens of MB/sec with no way to add capacity. You'd trade the entire parallelism story for a guarantee almost nobody needs. What people usually want is per-entity ordering, and that's free: key by user or account ID and every record for that key lands in one partition, ordered. "Global ordering costs you the system; per-key ordering is free — which do you actually need?" is one of the highest-value sentences in this chapter.
NOT_LEADER_OR_FOLLOWER, refresh metadata, retry — a blip of seconds. Its partitions are then under-replicated while the cluster copies data to restore RF=3, which is real background I/O.UnderReplicatedPartitions and OfflinePartitionsCount must both be 0 — anything else means durability is degraded or data is unreachable right now. ActiveControllerCount summed across the cluster must be exactly 1: zero means nobody's in charge, two means split brain. ISR shrink/expand rate flags a broker flapping from a dying disk or GC.UnderReplicatedPartitions is back to 0 — taking out two of three replicas turns a routine upgrade into an outage. Place replicas rack- and AZ-aware so RF=3 means three failure domains, not three boxes on one switch.Probes here are checkable, so they come fast:
acks=all. Walk me through exactly how you can still lose data." — they want min.insync.replicas, ISR shrink, unclean election. The classic near-miss.| Decision | Chose | Over | The cost I accepted |
|---|---|---|---|
| Storage | Append-only segments + sparse index, page cache, no app cache | In-memory store; dense index | A short forward scan per read; cold reads of old data hit disk hard |
| Durability default | RF=3, acks=all, min.insync.replicas=2 | acks=1 | Extra round trip in the p99; writes rejected when two replicas are down |
| Unclean election | Off for business data, on for telemetry | One global setting | Business topics can go fully offline; telemetry topics can silently lose and diverge |
| Delivery | At-least-once + idempotent consumers | Transactions everywhere | Every consumer must be written idempotently — a discipline enforced by review |
| Ordering | Per-partition only, key by entity | Global ordering | No total order across a topic; cross-entity sequence rebuilt downstream |
| Consumer flow | Pull with long polling | Broker push | Slight latency floor from polling; consumers own their pacing |
| Retention | 7 days default, per-topic override | Retain everything | Retention drives cost — 130 brokers instead of 55; needs an owner |
A junior asks: "Kafka writes everything to disk and it's still faster than our in-memory queue. How?" Explain in five or six sentences.
Disks are only slow when you make them jump around, and a log never does — it only appends to the end of one file, and sequential writes run hundreds of times faster than random ones. Kafka also keeps no copy of messages in its own memory: it writes into the operating system's page cache, so a broker with 55 GB free taking 24 MB/sec has roughly the last half hour sitting in RAM when consumers ask for it. Because the broker never inspects those bytes, it hands them from page cache straight to the network card with sendfile, skipping two copies and never touching the application heap. Producers batch and compress many records per request, so per-message overhead nearly vanishes. Your in-memory queue is fighting garbage collection and redoing work the kernel would have done for free. The lesson is bigger than Kafka: match your access pattern to what the hardware is good at.
Requirement flip: your biggest topic must retain 90 days instead of 7, with no budget increase. Run the numbers and commit to a plan.
86 TB/day × 90 × RF3 = 23 PB, roughly 1,650 brokers, for zero extra throughput — and the NIC is still at 3%. Huge disk with an idle network is the tell that broker disks are the wrong tier. Plan: keep 2–3 days hot on brokers (enough for every real-time consumer plus an outage) and turn on tiered storage so older segments live in object storage. Broker count falls back toward throughput-driven sizing, and 90 days costs object-storage rates instead of NVMe rates. Named cost: old-segment reads are much slower and go over the network, so a consumer that falls days behind leaves the fast path entirely.
10x this: 10 million messages/sec, but average size drops to 200 bytes. What breaks first, and is the answer still 130 brokers?
Bytes only double, to 2 GB/sec, so disk stops being the binding constraint. What breaks is message count: request rate, per-record CPU, and partition count. At 200 bytes, batching decides everything — raise linger.ms and batch size or the fleet drowns in per-request overhead. Partitions must grow roughly 10x for consumer parallelism, and with ~1,750 partitions per broker as a practical ceiling, that alone can set broker count. So the shape flips: the cluster becomes CPU- and metadata-bound rather than disk-bound, and the structural move is splitting into several clusters by domain rather than growing one.
Requirement flip: every event for a given user must be strictly ordered — and one automated account produces 30% of all events. Design it.
Per-user ordering is free: partition by user_id. But the heavy account creates a hot partition (chapter 2.6) — 30% of 1 GB/sec on one broker's disk, far past the ~30 MB/sec plan, read by one consumer. Options: (a) give that account its own topic, clean but it leaks a business fact into topology and doesn't generalize to the next whale; (b) split the key for known-heavy accounts into user_id:bucket across a few partitions, sacrificing strict ordering for that account only; (c) keep one partition and make its consumer trivially fast. I'd commit to (b) — but the senior move first is asking whether that account's events actually need mutual ordering. Automated feeds usually don't, and the answer changes the design.
Estimation drill, 90 seconds: a group is 4 hours behind on a topic taking 200 MB/sec. How long to catch up, and what's the hidden cost?
Backlog = 200 MB/sec × 14,400 s = 2.88 TB. Consumers must outrun the 200 MB/sec still arriving; assume they sustain 500 MB/sec, leaving 300 MB/sec of surplus. 2.88 TB ÷ 300 MB/sec ≈ 9,600 s ≈ 2.7 hours. The hidden cost is the point: 4 hours is far past the ~38-minute page-cache window, so every catch-up read comes off disk and competes with the write path — producer latency rises and other consumers slow while this group heals. Catch-up is a cluster-wide event. Mitigations: a consumer quota, or a business decision to skip ahead if the data is disposable. Asking for that second option out loud is the senior move.
sendfile beat "keep it in memory," and RAM ÷ ingest rate gives you both the performance model and the real reason lag alarms matter.acks=all alone is not durable because the ISR can shrink to the leader, so min.insync.replicas=2 is what turns a false promise into an honest error, and unclean election is an availability-versus-loss switch set per topic class.Every design in this book ends the same way. "I'd watch p99 latency, queue depth, error rate. I'd alert on the burn rate." By now you've promised a dashboard a hundred times. Today the interviewer collects the debt: "Design the dashboard. Design a metrics and monitoring platform — something like Datadog or Prometheus at cloud scale."
Feel the stakes before you draw a box. It's 3 a.m. and a payments database is down. The on-call engineer opens the dashboard to see when the errors started — and the dashboard is a spinner. The monitoring system shared a network with the thing it was monitoring, and they died together. Now the company is debugging an outage with ssh and hope. A metrics platform is the one system whose availability matters most at the exact moment everything else is failing. That constraint will follow us through the whole design.
And there's a second trap hiding in the prompt. This is a write-heavy firehose — millions of points per second, arriving forever, whether or not anyone is looking. Most systems in this book were read-heavy; this one inverts the ratio. Reads are rare, bursty, human. Writes never stop. And buried inside the write path is the real crux, the thing that has taken down more monitoring systems than any hardware failure: one engineer, one innocent-looking label, and an explosion you'll see coming by minute nine of the math.
Today's prompt is the dashboard you have promised at the end of every design in this book — a metrics and monitoring platform at Datadog or Prometheus scale. Scope it out loud: "I'll build metric ingestion from customer hosts and services, dashboard queries, and alerting. Metrics only — logs and distributed tracing are separate products, and I'll actually use that separation as a design tool later. Out of scope: billing, the UI itself, anomaly detection. Fair?"
Functional: agents on customer machines push numeric measurements with names and labels; users query them for dashboards ("p99 latency for checkout, by region, last 6 hours"); users define alert rules that page a human when a condition holds.
Non-functional, with numbers committed:
That last line is a non-functional requirement most candidates never state, and it's the one interviewers at any company that has lived through an outage are listening for.
Do the one multiplication that shapes everything:
Say the conclusions: "Five million writes a second means the write path is the product; I'll design it first and make it boring and horizontal. Seven terabytes a day means I need roughly a 10x compression win and a downsampling story or the economics fail. And fifty million series means my true scaling axis isn't points per second — it's series count. Points are cheap; series are expensive. That's where I'll spend my deepest dive."
The unit of everything is a series: one metric name plus one exact combination of label key-values. http_request_ms{service="checkout", endpoint="/pay", status="500"} is one series. Change any label value and it's a different series with its own stream of points. This definition looks like trivia. It is actually the bomb in the basement — hold that thought.
POST /api/v1/ingest — the agent endpoint. Batched and compressed: hundreds of points per request, metric names and label sets sent once per batch and referenced by index, API-key auth per tenant. At 5M points/sec, per-point HTTP requests would be absurd; batching is the first 100x saving and it happens before our servers see a byte.GET /api/v1/query?expr=p99(http_request_ms{service="checkout"}) by (region)&from=-6h&step=60s — a small expression language: select series by label match, aggregate, group, over a time range.POST /api/v1/alerts — a rule is a query expression, a threshold, and a duration: "fire if this holds for 5 minutes."Storage model, two parts. Points: (series_id, timestamp, value), append-only, time-ordered. Index: an inverted index from each label pair to the set of series ids that carry it — service="checkout" → {17, 40982, …} — so a query with three label matchers is three posting-list intersections, exactly like a search engine (chapter 3.19 will build the full-size version). Queries never scan points to find series; the index finds series, then points are read by id and time range.
Write path. The agent on each host collects measurements, batches 10 seconds' worth, compresses, and pushes. It also keeps a small disk-backed buffer — minutes of data — which is our degraded mode when intake is unreachable (that's the self-dependency section). The intake fleet is stateless: authenticate the tenant, validate the payload, enforce cardinality budgets (crux, coming), then produce to Kafka. Stateless means an intake node is disposable and the fleet scales linearly with load — at 5M points/sec I want zero cleverness here.
Why Kafka in the middle? Chapter 1.11's rule: put a queue where the producer's health must not depend on the consumer's. Agents must never be told "storage is slow, stop measuring" — during an incident, write volume spikes at the precise moment storage is struggling. Kafka absorbs the spike, and an hour of buffered stream (a few hundred MB/s batched and compressed — under a terabyte) means a storage-layer incident costs us delay, not data. Partitions are keyed by hash(series_id), so every point for a given series lands in the same partition, in order — which is exactly what the next stage needs.
Stream aggregators consume the partitions and do the 10-second windowing work (deep dive 1), then write to TSDB shards — the time-series database, sharded by series id, each shard replicated ×2. Because points for one series always flow through one partition to one shard, a shard owns its series completely: no cross-shard writes, ever.
Read path. A dashboard query hits the query service, which uses the index to resolve label matchers to series ids, fans out to only the shards owning those series, and each shard returns partial aggregates — not raw points — which the query service merges. Alert evaluation branches off the aggregator stream instead of querying at all. Both of those sentences hide the same idea, and it's deep dive 1.
Raw points at 10-second resolution are what agents send. But almost every question humans ask is coarser: "requests per minute," "p99 over 5 minutes." So the aggregators fold each series' raw points into 10-second windows storing min, max, sum, and count. Those four numbers are magic for one reason: they are mergeable. Sum of sums is the sum. Min of mins is the min. Count of counts is the count. Merge windows across time and you get 1-minute and 1-hour rollups for free; merge across hosts and you get the fleet view; merge across shards at query time and fan-out works. Any average is sum ÷ count, exact, at every zoom level. This is why shards can return partial aggregates instead of shipping millions of raw points to the query node — a 10–100x reduction in bytes moved and stored, and it's the same partial-aggregate trick chapter 2.9 used in stream jobs.
Then someone asks for p99, and the magic dies. Percentiles do not merge. You cannot combine two p99s into anything meaningful — not with an average, not with a max, not with anything. Here's the two-host proof, worth performing on the whiteboard:
So what do you store, if not percentiles? You store a sketch: a small, lossy summary of the whole distribution — think of it as a histogram with cleverly chosen buckets. t-digest keeps buckets finer near the tails, where percentiles live; DDSketch uses exponentially growing buckets, which guarantees the answer is within a fixed relative error, say 1% — your p99 of 2,000 ms might read 1,985 or 2,015, never 1,025. These are new names but an old friend: chapter 2.10's trade, bounded error for tiny constant memory — a few KB summarizing millions of values. And the property that makes them the answer here, the actual reason the industry settled on them, is that sketches merge. Combine host A's sketch with host B's and you get exactly the sketch of the combined traffic; merge sketches across 10-second windows and you get the hour. So the aggregator's window record is really: min, max, sum, count, and a sketch. Every percentile at every zoom level, from mergeable parts.
Two schools each report their 99th-percentile exam score. Averaging those two numbers tells you nothing about the combined 99th percentile — it depends entirely on how many students each school has and the shape of everyone's marks. But if each school hands over a compressed tally — "so many students scored 40–50, so many 50–60…" — you just add the tallies bucket by bucket and read off any percentile of the combined student body. The tally is the sketch. The analogy's one crack: real sketches blur bucket edges slightly to stay small, which is why the answer is "within 1%," not exact.
Everything in this platform scales with the number of series, not the number of points. A series is cheap to sample but expensive to exist: an index entry, an in-memory head chunk on a shard, a row in every fan-out. And series count is set not by us but by our users' label choices, multiplicatively. One engineer adding one label with unbounded values — user_id, request_id, session — can mint more series in an hour than the platform accumulated in a year. The write firehose is predictable; cardinality is the axis that explodes without warning. Defending it at intake, before bad series are born, is the difference between a monitoring platform and a beautiful outage machine. This is where I'd spend my remaining depth, and where the interviewer is hoping I'll go.
Walk the multiplication. http_request_ms with labels service (50 values), endpoint (40), status_class (5): at worst 50 × 40 × 5 = 10,000 series. In practice fewer — only combinations that actually occur are born — but bounded, because every label is bounded. Add host across 1,000 machines and you're at a few million series. Big, budgeted, fine.
Now a well-meaning engineer, debugging one slow customer, ships http_request_ms{…, user_id="…"}. Five million active users. Every unique combination that occurs becomes a brand-new series: 5M users × even a handful of endpoint combinations each ≈ 50 million new series — the entire platform's existing total, doubled by one line of code, in one afternoon. Each new series wants ~1–2 KB of head memory and index space on a shard: 50–100 GB of RAM demanded from machines provisioned for steady state. Index posting lists bloat; every dashboard query touching that metric now merges millions of series; shards OOM one by one. And tomorrow those users are back with new session labels — dead series pile up in the index like unclaimed mail. This exact failure is so common in the Prometheus world that the official documentation warns in bold that every unique key-value label combination is a new time series, and that unbounded values like user IDs and email addresses must never be labels; community postmortems of Prometheus servers OOMing from a single high-cardinality label are a genre of their own.
A sorting office builds one pigeonhole per unique address pattern it has ever seen. Sort mail by city × neighborhood × street and you need a big wall of pigeonholes — but a bounded wall. Let one clerk add "× recipient's full name" to the sorting scheme and the office now needs a pigeonhole per person: the building fills with millions of tiny boxes, nearly all holding a single letter, and finding anything means walking past all of them. The letters (points) were never the problem — the boxes (series) are. And the fix is the same in both worlds: the rule about what counts as a valid sorting key gets enforced at the front door, not discovered in the warehouse.
Defenses, in the order they should fire:
Inside a shard, storage is LSM-flavored, tuned for time. Incoming points append to a write-ahead log (crash recovery) and an in-memory head block holding each series' most recent chunk — recent data is both the hottest to query and the cheapest to serve from RAM. Every two hours the head is sealed into an immutable, time-partitioned block on disk with its own inverted index; background compaction merges small blocks into bigger ones. Deleting expired data is dropping whole blocks — no tombstones, no surgery. Append, seal, compact: the LSM idea from chapter 1.13, with time as the natural partition key.
Compression is where this design's economics live, and the intuition takes two sentences. Timestamps arrive on a metronome — every 10 seconds — so store the change of the change (delta-of-delta), which is almost always exactly zero: one bit per timestamp. Values barely move between neighbors — a temperature gauge reads 42.1 then 42.1 — so XOR each value's bits against the previous value's and store only the handful of bits that differ, which for an unchanged value is again a single bit.
Both tricks come from one paper. By 2013 Facebook's monitoring store — ODS, backed by HBase — had become too slow to be useful during incidents: a query over a few thousand series took tens of seconds, at exactly the moment seconds mattered. The team's observation was that virtually all debugging queries touch only recent data. So they built Gorilla, an in-memory database holding the most recent 26 hours as a write-through cache in front of HBase, and published it at VLDB 2015. The compression results made the paper famous: with delta-of-delta timestamps, 96% of timestamps compressed to a single bit; with XOR value encoding, about 51% of values did too; a 16-byte point shrank to 1.37 bytes on average — a 12x win that made "26 hours of everything, in RAM" affordable at an ingest rate of 700 million points per minute. Query latency dropped over 70x, to a 90th percentile of 10 milliseconds — and the behavior change is the part worth remembering. ODS was serving 450 queries per second when Gorilla launched; Gorilla soon handled over 5,000 steady-state, peaking near 40,000, because engineers built analysis tools nobody would attempt against a database that answered in tens of seconds. Make reads cheap and people ask more questions. The paper's encodings became the industry default; Prometheus's storage engine uses Gorilla-style XOR compression directly. When you write "delta-of-delta + XOR" on the whiteboard, you're citing one of the most-adopted systems papers of its decade.
Compression makes raw retention affordable for days; downsampling makes 13 months affordable. Three tiers, each an honest statement about what questions old data answers:
The query planner picks the tier by time range — "last 3 hours" reads hot, "last quarter" reads cold — and can stitch tiers at the boundary of a long range. Users never ask; the system chooses. Downsampling jobs are idempotent (rerunning a window produces identical output — chapter 2.2's habit), so a crashed job is re-run, not repaired.
The obvious alerting design is a poller: every 30 seconds, run each rule's query against the TSDB. The estimate already killed it — 100K rules is 3,300 heavyweight queries per second of pure background load, most of them re-reading the same freshly written data over and over, competing with the humans mid-incident who actually need the query path. Polling is a design that works in the demo and bankrupts you at scale.
Instead, evaluation moves into the stream. Alert rules are compiled and pushed down to evaluator nodes that consume the same aggregated windows flowing out of the aggregators; each evaluator holds its rules' state in memory — a small sliding window per rule — and updates it incrementally as data arrives. Evaluating 100K rules becomes filtering a stream you were already producing: no TSDB load at all, and alert latency equals ingest freshness, ~30 seconds. The honest footnote: gnarly rules — long lookbacks, joins across many metrics — still fall back to scheduled queries. Streaming for the 95%, polling for the awkward tail.
Then make the alerts humane, because a noisy pager is a disabled pager. A rule fires only after its condition holds for a configured duration ("for 5 minutes"), which forgives one bad scrape. It clears at a lower threshold than it fires (hysteresis) so a metric oscillating around the line doesn't page in a loop, and identical firings dedupe into one incident. And the platform natively supports burn-rate alerts on SLOs — chapter 2.8's fast-window-plus-slow-window pattern — because "the error budget is burning 14x too fast" is the page that matters, and customers shouldn't have to hand-build it.
Now the requirement I flagged in the scoping list. A monitoring platform that shares fate with what it monitors is a smoke detector wired to the house mains: it dies in exactly the fire it exists to announce. So the platform runs in its own failure domain — separate accounts, separate clusters, ideally separate regions from the monitored infrastructure; it must never depend on the customer's databases, queues, or service mesh. Internally the same razor applies one level down: the component that delivers pages depends on almost nothing — if every TSDB shard is dead, "evaluate stream, send page" must still work, and the final notification hop rides an independent third-party path, because a pager that depends on the platform it reports on has the same disease one layer deeper.
The agent is the last line. When intake is unreachable, agents keep collecting and spool to their local disk buffer — say 30 minutes — because the customer's incident review will want the data measured during our outage. On reconnect, agents send the freshest points first (alerting needs now, not twenty minutes ago) and trickle the backlog behind them, rate-limited by intake so that 100K agents flushing simultaneously don't become the retry stampede of chapter 2.7 and kill the intake that just recovered.
The monitoring vendor's nightmare happened in public. At 06:03 UTC, a routine security update of systemd rolled out automatically — through a legacy unattended-upgrades channel nobody remembered was enabled — across tens of thousands of Datadog's virtual machines in five regions, spread over three separate cloud providers. On each updated host, restarting systemd-networkd deleted the network routes that the container networking plugin (Cilium) had installed, so the nodes fell off the network within the same hour. Datadog's web app, dashboards, and monitors degraded for customers worldwide; full recovery took about two days, and the company published an unusually detailed postmortem. Three lessons map straight onto this chapter. First: the trigger wasn't customer traffic — it was one change applied everywhere at once; the remediations included disabling that legacy update channel and moving base-OS changes to staged, region-by-region rollout, like any other deploy. Second: during recovery Datadog explicitly restored live telemetry and alerting before backfilling history — freshest-first, exactly the ordering our agents implement, and the same lesson Gorilla's authors listed first. Third: customers whose own systems were perfectly healthy were blind anyway, because their eyes lived in someone else's failure domain — the self-dependency rule doesn't stop applying just because monitoring is a vendor's problem. Some teams now keep a minimal independent probe — even a crude external ping — precisely so that "is the monitoring down, or is everything down?" has an answer.
Intake overload. A tenant's deploy goes wrong and their agents triple output; or a cardinality bomb slips past budgets. Intake sheds by priority, chapter 2.7 style: platform-health metrics and any series referenced by an alert rule are protected; unreferenced custom metrics shed first, per-tenant, with the drop counted and visible to that tenant. Degrade the nice-to-haves, defend the pager.
TSDB shard death. Two replicas per shard; a dead replica's partner serves reads while a new one rebuilds by replaying Kafka — the buffer doubles as the recovery log. If a time range is missing from all replicas, the query API must say so: partial data gets a partial flag, rendered as a gap. Never zero-fill — a silent gap reads as "the metric dropped to zero," which pages someone for a ghost, or worse, convinces an absence-based alert that everything is fine. In monitoring, honest gaps beat smooth lies; a wrong chart is worse than no chart.
Aggregator crash. Windows in flight are lost — but the source points are still in Kafka, so the replacement replays from the last committed offset and rebuilds the same windows. Kafka's retention is our exactly-once eraser: reprocessing produces identical windows (idempotent again), so the failure costs freshness, never correctness.
What I'd watch — the platform's own golden signals: end-to-end ingest lag (measurement time to queryable, the freshness SLO), dropped-point rate by cause (budget, shed, malformed — each cause is a different conversation with a different tenant), per-tenant series growth rate (the fuse), Kafka consumer lag per partition, alert-evaluation delay, and query p99 by tier. Rollouts: the platform deploys like any serious system — canary one intake node, one shard replica, one evaluator; watch its own metrics; auto-roll-back on freshness regression. The Datadog incident is the standing reminder that "apply everywhere at once" is an outage pattern, not an efficiency.
At 10x — a million hosts, 50M points/sec, half a billion series: intake and Kafka scale linearly (stateless fleet, more partitions). What breaks first is series-shaped, as the crux predicted: index memory per shard and query fan-out width. The moves: more shards via consistent hashing with shard splitting (chapter 1.9), a two-level query tree so the root merges dozens of partial merges instead of thousands of shard responses, and per-tenant cell isolation (chapter 2.5) so one giant tenant's cardinality accident is contained inside its own cell instead of degrading a thousand neighbors.
What the interviewer probes next, nearly verbatim:
user_id label at 2 p.m. Walk me through the next ten minutes — what breaks first, and where do you stop it?" (They want the explosion math, then intake budgets — defense before the shards, not cleanup after.)| Choice | Why | What it costs |
|---|---|---|
| Kafka buffer between intake and storage | Absorbs incident-hour spikes; storage trouble costs delay, not data; doubles as shard-rebuild log | ~30s added freshness; a large cluster to operate; one more system that must outlive outages |
| Pre-aggregate into 10s windows of min/max/sum/count + sketch | Mergeable across hosts, windows, shards; 10–100x storage and query-bytes win | Sub-10-second spikes are invisible; raw sample values are gone forever |
| Sketches for percentiles (t-digest / DDSketch) | Percentiles don't merge; sketches do, with ~1% bounded error in KBs | Approximate answers; explaining "why is p99 slightly different today" to exact-minded customers |
| Series budgets + rejection at intake | Cardinality bombs contained before shards allocate anything | Deliberately drops customer data; budget-raise requests become a support workflow |
| Streaming alert evaluation | 100K rules for free on a stream you already produce; alert latency = ingest freshness | Evaluator state to manage; complex rules need a query-based fallback path anyway |
| Three retention tiers with downsampling | 13 months affordable; queries pick resolution by range automatically | Old data loses resolution permanently; downsampling jobs are one more thing to run and watch |
| Separate failure domain + agent disk buffers | Monitoring outlives the outages it narrates; measured-during-outage data survives | Duplicated infrastructure, real money; agents get heavier and need their own hygiene |
Explain to a junior engineer, in five or six sentences, how adding one label to one metric can take down an entire metrics platform.
A metrics platform doesn't store "a metric" — it stores one time series per unique combination of a metric's label values, and everything expensive (memory, index entries, query work) scales with how many series exist. Bounded labels like status code multiply into a bounded number of series, and that's fine. But a label like user_id has millions of values, so every user who shows up mints a brand-new series — one line of code can double the platform's total series count in an afternoon. Each new series demands index space and RAM on a storage shard, so shards run out of memory and queries that touch the metric slow to a crawl. That's why platforms enforce series budgets at the front door and reject or strip the label before storage ever sees it. The rule of thumb: if a label's values are unbounded, it belongs in logs or traces, not in a metric.
Rerun the sizing for an IoT fleet: 1 million devices, 20 metrics each, one point per 60 seconds. Points per second? Active series? Which parts of this chapter's design would you actually change?
1M × 20 / 60s ≈ 333K points/sec — 15x less write volume than our design. But active series = 20 million, and device fleets churn: devices die and are replaced, so series accumulate. Cardinality stays a first-class concern while the firehose shrinks. You'd keep intake budgets, the index design, and retention tiers; you could shrink Kafka and the aggregator fleet dramatically, and might merge intake and aggregation into one stage. The lesson is the chapter's thesis inverted: points/sec dropped 15x but the design barely simplifies, because series count — not write rate — was always the load-bearing number.
Product flips a requirement: "customers must be able to see metrics per user — that's the feature." You can't just say no anymore. Design your way out without breaking the platform.
Route around metrics, don't bend them. Per-user questions are event queries, so ship the data as structured events into the logs/traces product, where each event carries user_id as an attribute and cost scales with event volume, not distinct users. If the product needs per-user aggregates (say, a customer-facing usage graph), run a dedicated stream job keyed by user into its own store — a deliberate, provisioned high-cardinality system with its own budget, not a label smuggled into the shared TSDB. What reads senior: naming the boundary explicitly — bounded dimensions are metrics, unbounded identities are events — and meeting the requirement by choosing the right system rather than heroically misusing the one on the whiteboard.
Design the "dead man's switch": customers want an alert when a metric stops arriving (a host died, an agent broke). Detecting absence across 50 million series sounds like scanning 50 million last-seen timestamps every 30 seconds. Make it cheap.
Scope, don't scan. Absence only matters where someone asked: track last-seen only for series referenced by an absence-type alert rule — thousands, not millions — as a tiny state map inside the streaming evaluators, which see every arriving window anyway; a timer wheel fires when a tracked series goes quiet past its threshold. For whole-host death, don't watch 500 series per host — watch one designated heartbeat series per agent. Second-order senior point: absence detection is exactly why gap honesty matters — if shard failures were silently zero-filled, a dead host would look alive at value zero, and the dead man's switch would sleep through the funeral.
A teammate proposes: "Storage is cheap now. Skip downsampling — keep raw 10-second data for the full 13 months. Simpler system, fewer jobs." Argue it properly, with numbers, before you disagree — or agree.
Steelman first: it deletes the downsampling jobs, the warm/cold write paths, and the tier-picking planner — real complexity, gone. Now the numbers: 7 TB/day raw is ~2.5 PB/year uncompressed; Gorilla-style compression (~12x) brings it to ~200 TB — storable, and in object storage even affordable. But storage was never the real bill: a "last quarter" dashboard query now reads ~800K raw points per series instead of ~2K hourly rollups — a 400x penalty on exactly the wide, slow queries humans run while impatient. You'd end up pre-computing coarse aggregates at query time, repeatedly, which is downsampling with worse economics. Verdict: keep tiers; concede that if the product were "raw forensic archive, rarely queried," the teammate wins — the design follows the query pattern, not the storage price.
What breaks first at 10x (one million hosts, 50M points/sec, 500M series) — and what's your first structural move?
Not intake, not Kafka — both scale linearly by adding stateless nodes and partitions. Series-shaped things break first: shard head/index memory (500M series × ~2 KB ≈ a terabyte of RAM to distribute) and query fan-out, where the root now merges partials from hundreds of shards and the merge itself becomes the bottleneck. First moves: split shards along the consistent-hash ring, and add an intermediate merge layer so the query tree is two levels deep. The structural move worth naming: per-tenant cells — a giant tenant gets its own intake-to-TSDB slice, so its cardinality accidents and query storms are contained. If you said "add more Kafka," you scaled the part that was already fine.
The interviewer writes one line: "Design search for our marketplace. Five hundred million listings. Keyword queries, ranked by relevance, fast."
You say "I'll use Elasticsearch." They smile and ask the question this chapter exists for: "Sure. Now open it up. What's inside, and why is it shaped like that?"
Here's the pain that makes the shape inevitable. Your listings live in Postgres. A user types waterproof hiking boots. The naive query is WHERE description LIKE '%waterproof%', and Postgres must read all 500 million rows, because a B-tree answers "starts with" and never "contains" (Part 1.3). That's ~500 GB of text off disk per query. Minutes, not milliseconds. And if it did finish, it would hand back 3 million rows in primary-key order — an order with nothing to do with which boots the user wants. Two problems, not one: finding is slow, and ranking doesn't exist.
The fix is to stop searching the documents and search a structure built from them. That structure is the inverted index. Building one that's sharded across machines, scored in milliseconds, and still fresh seconds after a seller edits a price — that's the whole interview.
Functional, agreed out loud:
"trail running" matches those words adjacent, in order.Out of scope, with reasons: typeahead in the search box is a different machine with a different traffic profile — ten requests per search instead of one — and gets its own design (3.2). Semantic retrieval I'll gesture at as a reranking stage and leave to 3.24.
Non-functional, committed:
Search houses two tenants: queries, which want sorted, unchanging structures held in memory, and indexing, which wants to mutate them constantly. Every hard decision below comes from that fight.
Index size. 500M × 1 KB ≈ 500 GB of text, ~150 words per doc, ~100 distinct terms after normalization. That's 500M × 100 = 50 billion postings — 50 billion "term X appears in doc Y" facts. Stored naively that's over a terabyte; stored the way real engines do it — doc IDs sorted, delta-encoded, bit-packed — the classic IR rule holds: an index with positions lands at 30–50% of the source text, without positions 10–15%. Phrase queries are a requirement, so I'm paying for positions. Call it 250 GB.
That doesn't fit one machine's page cache, and one machine is one failure. Shards want to be 20–40 GB — small enough to merge, move, and restore without drama. So 12 primary shards — about 21 GB each today, with room to double before they leave that range and force a resharding (which means a full reindex, so headroom is cheap insurance) — each with 2 replicas.
Query CPU sizes the fleet. 10,000 QPS × 12 shards = 120,000 shard-queries/sec. At 4 ms of CPU per shard-query that's 480 CPU-seconds per second — ~500 cores, ~16 machines at 32 cores. Storage wanted ~18. They agree: 20 nodes with headroom. Now read the sensitivity, because it turns index internals from trivia into architecture:
At 120,000 shard-queries per second, every extra millisecond of per-shard work costs about 120 cores.
Skip pointers, compressed posting lists, and a cheap first-pass scorer are each worth a rack.
Indexing. 5,000 docs/sec × 1 KB = 5 MB/s of new text; after merge write-amplification (each document's data gets rewritten three to five times over its life) call it 15–25 MB/s of disk writes. Trivial on SSD. Steady state is cheap. Bursts aren't: a full rebuild of 500M docs at 50,000/sec takes about 3 hours. Hold that number — it's our disaster-recovery time, our migration window, and the reason one decision later goes a particular way.
GET /v1/search?q=waterproof+hiking+boots&category=footwear
&price_max=15000&from=0&size=10
200 OK
{ "total": {"value": 10000, "relation": "gte"}, // approximate past 10k
"partial": false, // see failure modes
"hits": [{"id": "L-9931", "score": 14.2, "snippet": "...boots..."}],
"facets": {"brand": [{"k": "Salomon", "n": 812}]} }
Two honesty flags in the contract: total is a bound, not a count (counting all matches costs as much as the search), and partial tells the caller a shard didn't answer. Both pay off later.
The indexed document is denormalized — search can't join:
{ id, title, description, // analyzed text, positions on
category, brand, condition, // exact-match keyword fields
price_cents, rating, created_at, // numeric: filter, sort, facet
seller_id } // NOT seller_rating — see below
Note what's missing. Seller rating is tempting to denormalize, until you notice one rating change rewrites all 200,000 of that seller's listings. High-churn fields do not belong in an inverted index. I keep seller_id and look the rating up at rerank time. Stock and flash-sale prices get the same treatment: filter loosely in the index, verify against Postgres at checkout.
Write path. The app writes only to Postgres. Change data capture (CDC) — a reader that tails the database's own replication log, so it sees every committed row change in commit order — feeds those changes into Kafka; an indexer consumes, denormalizes, and sends batched bulk requests, each document routed by hash(docID) % 12. I refuse dual-writes (app writes DB, then writes search) for the Part 1.14 reason: two writes with no shared transaction drift apart silently and forever. Kafka isn't decoration — it's the buffer that absorbs reindex bursts and the tape I replay when the indexer has a bad day.
Read path. Coordinator analyzes the query text with the same pipeline used at index time (a mismatch here is the most common relevance bug in production search), fans out to one replica of each shard, each scores locally and returns its top 10 as IDs and scores only; the coordinator merges 120 candidates into 10, then a second round fetches display fields for those 10. Two rounds on purpose — shipping full documents for 120 candidates would be megabytes per query.
The index at the back of a textbook. You don't find "sharding" in a 900-page book by reading from page one — you flip to the back, where someone inverted the book: for every important word, a sorted list of pages. That's an inverted index: instead of document → words, store word → documents. The list hanging off each word is its posting list. Now the honest part, and it's the whole chapter: a book index is built once, when the book is finished, and never changes. Ours must absorb five thousand edits a second while ten thousand people read it. Every awkward thing in this design comes from that one difference.
Tokenize and normalize. Split into terms, lowercase, strip accents so "café" finds "cafe". Easy in English, treacherous elsewhere: "C++" and "COVID-19" die under naive punctuation splitting, and Chinese and Japanese have no spaces, needing dictionary or n-gram tokenizers. Analyzer chosen per listing locale.
Stem — the recall/precision dial. Stemming chops words to a crude root, so "running", "runs", and "ran" become one term. It raises recall and lowers precision: the classic Porter stemmer maps both "university" and "universe" to "univers". My commit: light stemming, plus the unstemmed form in a second field, boosted at query time. The stemmed field finds it; the exact field ranks it above sloppier matches.
Stopwords — where old textbooks are wrong. Classic IR deleted "the", "a", "of" to shrink the index. I'm keeping them: storage is no longer the binding constraint, phrase queries collapse without them ("The Who", "to be or not to be"), and BM25's IDF already crushes their contribution to near zero. Deleting them hand-optimizes something the scoring function does better on its own — which is why Elasticsearch's default analyzer ships with an empty stopword list.
Synonyms — index time or query time? Index-time expansion ("tv" stored as both "tv" and "television") is fast at query time but bloats the index, skews term statistics, and — the killer — changing the list means reindexing 500 million documents. We estimated that at three hours. Query-time expansion rewrites the incoming query instead: slightly wider queries, but the list becomes a config change. I commit to query-time synonyms, because merchandising edits that list weekly from search logs and no weekly business decision should cost a three-hour rebuild. The earlier estimate decided this for me.
GitHub's old code search ran on Elasticsearch and struggled for years: users couldn't search for ++, or for an exact substring inside an identifier. The reason is everything above. A pipeline built for prose destroys code — it splits on punctuation that carries meaning, lowercases identifiers where case matters, stems things that aren't words. In 2023 GitHub shipped a from-scratch replacement, an engine written in Rust called Blackbird, and the central design choice was to abandon word tokenization for n-grams: index every short character sequence, so any substring query becomes an intersection of n-gram posting lists — the same trick Russ Cox had described for Google Code Search a decade earlier. At launch it covered roughly 15.5 billion documents across some 45 million repositories, with heavy deduplication because so much code on GitHub is forked and copied. The lesson isn't "use n-grams." It's that the analyzer is a product decision, not a default, and a world-class team can be beaten by its own tokenizer.
A query for hiking boots is an intersection. "boots" might be in 8 million documents, "hiking" in 2 million. Walking both lists end to end is 10 million steps for one term pair, 120,000 times a second. Dead.
Two properties save it. First, drive from the shortest list — the rarest term: 2 million checks, not 10 million. Second, skip pointers: every 128 entries, the posting list stores a checkpoint ("doc ID 4,410 lives at byte offset 96,300"). When the driver asks "does 'boots' contain doc 4,410?", you don't walk — you hop along checkpoints until you overshoot, then scan one small block. Linear becomes near-logarithmic, and the intersection finishes in single-digit milliseconds.
Skip pointers are the express train. The local stops at every station — that's walking a posting list one entry at a time. The express stops every 128 stations. To reach station 4,410 you ride the express until you overshoot, get off at the last stop before it, and walk the rest. You still cover the same line, but you touch a handful of stops instead of thousands. The analogy breaks in one useful place: the express and the local run on the same track here — the checkpoints are stored inside the posting list itself, so there's no second structure to keep in sync.
Sortedness earns its keep twice: it makes intersection a one-pass merge, and it makes the list compressible, because you store gaps instead of IDs (12, 7, 8, 17, …) and small numbers pack into a byte or two. That's how 50 billion postings fit in 250 GB and stay in page cache — which keeps per-shard work near 4 ms, which is how you avoid buying 120 cores per wasted millisecond. It all chains back to one property.
Finding 40,000 matching listings is the easy half. Putting the right ten first is the product.
The intuition: a term matters more if it appears often in this document (term frequency) and rarely in the corpus (inverse document frequency). In "gore-tex boots", "boots" is in 8 million listings and says almost nothing; "gore-tex" is in 5,000 and says nearly everything. That's also why we drive intersections from the rare term — the cheap plan and the correct plan agree.
Raw TF-IDF has two ugly failures, and BM25 — my committed default — fixes exactly those two. Saturation: under raw counts, a document with "waterproof" 60 times scores 30x one with it twice, which is absurd — the second mention is informative, the sixtieth means you're being spammed. BM25 bends the curve so early occurrences count heavily and later ones flatten. Length normalization: long documents win under raw counts just by being long, so BM25 divides by document length relative to the corpus average.
Concretely: doc A is a real product, 30 words, "waterproof" twice. Doc B is a keyword-stuffed 4,000-word page with it 60 times. Raw counts rank B far above A. BM25 saturates B's 60 mentions to barely more credit than A's two, then penalizes B for its length — and A wins, which is what a human would have said. That's the whole argument, and it's what I'd say at the whiteboard instead of writing the formula.
Say this out loud in the interview, because it's the slot every ranking model plugs into without touching retrieval. And raise its weakness before the interviewer does: phase 2 can only rerank what phase 1 retrieved. If BM25 never surfaces the listing because the user typed "sneakers" and the listing says "trainers", no model downstream can save it. That recall ceiling is exactly the hole vector retrieval fills, running a semantic candidate generator alongside the lexical one and fusing the lists (3.24). I'd add it as a second retriever, never a replacement — lexical retrieval is exact, debuggable, and cheap, and a user who types a part number expects that part number.
Everything that makes the index fast makes it hostile to updates. Posting lists are fast because they're sorted, compressed, and contiguous — so inserting one document into the middle means rewriting them. A single new listing touches ~100 terms, so one insert would rewrite 100 posting lists, some tens of megabytes long. At 5,000 documents per second that isn't slow, it's arithmetically impossible. And yet the requirement says a seller's edit must be searchable in seconds. The crux of search is reconciling an immutable, sorted read structure with a mutable, real-time write stream. The answer — segments — is the most elegant idea in this chapter, and it's why "near-real-time" has the word "near" in it.
The trick is to stop updating anything. New documents accumulate in a small in-memory buffer. On a fixed refresh interval — one second by default in Lucene-based engines — that buffer is frozen into a brand-new, tiny, fully-formed inverted index called a segment, written once and never modified. A shard isn't one index; it's a stack of segments, and a query runs against all of them and merges the results.
A library that refuses to re-shelve. The main catalogue is printed and bound, so nobody can slot a new card into the middle of it. Instead, every evening the librarian prints a thin supplement listing that day's arrivals. To find a book you check the main catalogue and every supplement, then combine what you found. A book that leaves the collection isn't erased from the paper — it goes on a "gone" list you check as you read. Once a month someone reprints the main catalogue with all the supplements folded in and the gone-list entries quietly dropped. Where the analogy stops being comfortable: paper supplements are free to ignore, but here every extra supplement is a real term lookup on every single query, so letting them pile up is how a fast index becomes a slow one.
Three consequences, all of which a senior candidate volunteers:
1. Freshness is quantized. A document is invisible until the next refresh. That's the "near": seconds by design, not a bug to apologize for. The interval is a dial — 1 second for the live seller-edit path, 30 seconds during bulk reindexing, because each refresh creates a segment and segments cost.
2. Segments must be merged. Every query consults every segment, so a shard with 500 tiny segments does 500 term-dictionary lookups per term; query cost grows with segment count. Background threads continuously merge small segments into big ones. The price is write amplification — each document's data rewritten three to five times over its life. You buy cheap writes and cheap reads by paying background I/O forever.
3. Deletes are lies. You can't remove a document from an immutable file, so a delete flips a bit in a per-segment tombstone bitmap. The data is still on disk, just skipped at query time; it vanishes only when a merge rewrites that segment. An update is a delete plus an insert. So a workload rewriting every document hourly doesn't update the index — it fills it with corpses and makes the merger thrash. Mitigation: the data-model rule from earlier, keep volatile fields out.
One durability decision falls out, and it's where I spend the "search is derived" freedom the first time. Lucene-based engines keep a write-ahead log (the translog) and can fsync it on every request. I'll relax that to fsync-per-batch, accepting that a hard crash loses the last second of indexing — because truth is in Postgres and the events are still in Kafka, so recovery is a replay, not a loss. In a primary datastore that trade would be reckless. In a derived one it's free performance.
Almost every search system you've touched traces to one person's side project. Doug Cutting wrote Lucene in Java in 1999 and put it on SourceForge; it moved to the Apache Software Foundation in 2001 and became a top-level project in 2005. Lucene isn't a server — it's a library implementing exactly this section: analyzer chain, term dictionary, compressed posting lists, immutable segments, background merges, tombstone bitmaps. Everything since has wrapped it. Solr, created by Yonik Seeley at CNET in 2004 and donated to Apache in 2006, wrapped it in an HTTP server. Elasticsearch came from Shay Banon, who had earlier written a Lucene wrapper called Compass and, by his own retelling, began the rewrite while building a recipe search app for his wife; the first release landed in February 2010 and added what Lucene never had — sharding, replication, and cluster coordination. So when an interviewer says "open up Elasticsearch," the honest answer is: inside is Lucene doing the segment dance on one machine, and around it is a distributed system whose whole job is to shard documents, fan out queries, and merge results. Two separate problems — and knowing which is which is most of this interview.
250 GB and 500 cores don't fit on one machine. There are exactly two ways to split, and one is a trap the interviewer wants you to walk into.
Shard by term (the "obvious" one, since the index is a term dictionary): terms A–F on shard 0, G–L on shard 1. Each posting list lives whole on one machine. Now run "hiking boots": "hiking" is on shard 1, "boots" on shard 0, and the intersection needs both lists in the same place. "boots" is 8 million entries — tens of megabytes on the wire, per query, at 10,000 QPS. You'd need terabits of internal bandwidth for a two-word query. Worse than bandwidth: popular terms make permanently hot shards (Part 2.6), and losing a shard removes a slice of the vocabulary — every search containing a G–L word simply stops working, which reads as broken, not degraded.
Shard by document (what everyone ships): each shard holds 1/12 of the documents and a complete index over just those. Every query goes to every shard, but each does its whole job locally with zero cross-talk and returns 10 IDs and 10 scores. Adding shards adds capacity linearly, and losing one costs 1/12 of the corpus — every query still returns mostly-right results instead of some queries returning nothing.
If doc-sharding feels like folklore, it isn't: it's in "Web Search for a Planet: The Google Cluster Architecture" (Barroso, Dean, and Hölzle, IEEE Micro, 2003), one of the quietly most influential systems papers ever written. It describes Google's index split into pieces each holding a randomly chosen subset of documents; a query is broadcast to one replica of every index shard, each returns its best local hits, results are merged, and a second pool of document servers fetches titles and snippets for the winners. Two phases, exactly as above. The paper's other argument was economic and equally durable: because the workload is embarrassingly parallel across doc shards, throughput comes from adding cheap commodity PCs rather than bigger machines, with replication supplying both throughput and fault tolerance. Twenty years on, the shard-and-merge shape of every search cluster you will ever operate descends from that paper.
Distributed IDF. BM25 needs to know how rare a term is across the whole corpus, but each shard sees only its slice. Because documents are routed by hash — effectively at random — each shard's local document frequency estimates the global one well and the scores merge acceptably. Say that out loud as an accepted approximation rather than hiding it. It breaks on tiny indexes and on non-random routing (route by seller or tenant and term distributions diverge wildly). The fix exists — a preliminary round gathering global term statistics, which Elasticsearch exposes as the DFS search type — at the cost of an extra round trip per query. My commit: approximate IDF in production, global stats only in the offline relevance-evaluation harness where correctness beats latency.
Tail latency compounds. With a fan-out of 12, a query is only as fast as its slowest shard, so the coordinator's p99 is roughly the shards' p99.9. One node in a GC pause or a heavy merge damages every query, not one in twelve. This is Dean and Barroso's "tail at scale", and the mitigation is hedging: if a shard hasn't answered by p95, send the same request to a second replica and take the first response — a few percent extra load to erase the tail.
Deep pagination. To return results 1,000–1,010 ranked globally, every shard must return its top 1,010, because nobody knows which shard holds the global #1,004. At page 10,000 that's a memory incident. I'll cap pagination near page 50 (nobody clicks past it — measure and prove it) and offer a cursor-based search_after API for exports that genuinely walk the whole result set.
An inverted index answers "which docs contain X" beautifully and "what's the category of each of these 40,000 docs" terribly — it's built the wrong way round for that. So engines keep a column-oriented structure beside the postings: doc values, one packed array per field ordered by document. Faceting becomes "walk the matching doc set, read one column, tally ordinals" — cache-friendly and fast; sorting by price uses the same structure. Filters are cacheable bitsets that AND cheaply and, contributing no score, get reused across queries. I enable doc values only on fields we facet or sort on, since otherwise they're pure added storage. One honesty note: facet counts on a distributed index are approximate for high-cardinality fields, because each shard returns only its own top buckets — which is why real APIs return an error bound beside the count.
A shard has no live replica. Fan-out returns 11 of 12 answers. Two choices, and I want this decided explicitly rather than by a default. For marketplace search I commit to partial results with the flag set: search is already an approximation and an empty page loses a sale — but the flag is returned, logged, and alerted on, never silently swallowed. I'd flip that instantly for another product: legal e-discovery or a compliance audit must fail loudly, because "no results" and "we didn't look at 8% of the corpus" are different answers and only one is honest.
The index is corrupt, or an analyzer change must ship. Same answer, and it's the payoff of the derived-store rule: rebuild from source. You cannot change tokenization in place — existing segments hold terms from the old analyzer, and mixing them causes silent relevance rot. So: new index with the new mapping, replay from Kafka and backfill from Postgres, shadow queries comparing top-10 overlap, then flip an alias atomically so readers switch in one step, keeping the old index a day as rollback (Part 2.11, verbatim). That rebuild takes ~3 hours by our estimate, which is why we also snapshot to object storage hourly: restoring 250 GB is minutes, with the full reindex as the guaranteed-correct fallback behind it.
The dashboard, in the order I'd read it during an incident:
What breaks at 10x (5B docs, 100K QPS), in order: first the fan-out, since 10x the corpus means ~120 shards and a p99 that's the worst of 120 — forcing routing, where locale or category joins the shard key so a query touches a subset, plus hedging for the rest. Second, merge I/O, as write amplification on 2.5 TB per replica saturates disks; the fix is separating indexing from serving — build segments on dedicated indexing nodes and ship the finished immutable files to serving nodes, which is exactly what immutability makes possible. Third, deep pagination and faceting, which stop being expensive and start being outages. None of these is "the inverted index doesn't scale." The index is fine; the fan-out and the merge pipeline bend.
A passing senior answer names doc-sharding and defends it against term-sharding unprompted, says "immutable segments" and explains why deletes are tombstones, commits to BM25 with the saturation-and-length argument instead of the formula, and volunteers that freshness is seconds and why. Then the probes come, usually these:
| Decision | Why | What it costs |
|---|---|---|
Inverted index, not LIKE | A 500 GB scan becomes a few list intersections, and ranking comes along free | A second copy (~250 GB) that must be kept in sync forever |
| Shard by document (12 shards) | Self-contained shards, linear scaling; failure loses documents, not vocabulary | Every query touches every shard: tails compound, IDF goes approximate |
| Immutable segments + merges | The only way to get real-time writes on a sorted, compressed structure | 3–5x write amplification, permanent background I/O, deletes linger until merge |
| 1s refresh (30s during bulk) | Meets the 5-second freshness target with room | Freshness quantized, and more segments means slower queries |
| BM25 first, model second | Cheap over millions, expensive over 1,000; ML plugs in without touching retrieval | Recall ceiling: phase 2 can't rank what phase 1 never found |
| Query-time synonyms | The list changes weekly; a 3-hour reindex per edit is unacceptable | Wider queries at runtime; multi-word synonyms behave badly |
| Keep stopwords, light stemming + exact field | Phrase queries need stopwords, IDF de-weights them anyway, exact field restores precision | Bigger index and two fields to keep in step |
| CDC from Postgres; index derived | No dual-write drift, always rebuildable, lets me relax fsync durability | Seconds of lag, a pipeline to operate, no read-your-own-writes |
| Partial results on shard loss | An empty results page loses the sale; search is already approximate | Users can silently get worse results — flag, log, and alert are mandatory |
Explain to a junior engineer, in five or six sentences, what an inverted index is and why search engines rebuild it out of small immutable files instead of updating it in place.
An inverted index is the back-of-the-book index for your whole dataset: instead of document → words, you store word → the sorted list of documents containing it, so searching becomes "fetch two short lists and intersect them" instead of "read every document." Keeping those doc-ID lists sorted is what makes it fast — sorted lists intersect in one pass, compress to a byte or two per entry, and let you skip ahead using checkpoints instead of walking. But that same sortedness means inserting one document would rewrite the middle of a hundred compressed lists, far too expensive to do thousands of times a second. So engines never edit the index: new documents pile up in memory and, once a second, get frozen into a small brand-new index file called a segment, and a search runs against all the segments and merges the answers. Background threads combine small segments into bigger ones so the count stays manageable, and deleting a document just flips a bit in a "still alive" bitmap until a merge finally drops it. That's why search is near-real-time: your update appears on the next refresh, a second later, not the instant you save.
Now add exact-phrase search. A user searches "trail running shoes" in quotes. Your postings store doc IDs and term frequencies. What must you add, what does it cost, and how does the query execute?
Positions: each posting becomes doc ID, frequency, and the token offsets where the term appears. Execution is two-step: intersect the three posting lists for candidate documents, then for each candidate walk the position lists together, keeping only documents where a position of "trail" is followed by "running" at +1 and "shoes" at +2. Cost: positions roughly double or triple the index (they're why we budgeted 30–50% of corpus size rather than 10–15%), and per-candidate verification makes phrase queries meaningfully slower than bag-of-words ones. Two things a senior adds: this is why you can't strip stopwords if you want phrase search ("The Who" needs "the" and its position), and if phrase queries are rare you can keep positions on the title field only and accept degraded phrase matching in the body.
10x the write side only. Corpus stays at 500M documents, but a pricing engine rewrites every listing's price every 10 minutes: ~830,000 updates/sec. What breaks, and what do you do?
Everything, and the failure is structural, not capacity. Every update is a delete plus an insert, so the index accumulates 830K tombstoned documents per second; merges must continuously rewrite the entire 250 GB just to reclaim dead space, saturating disk I/O; segment counts climb, queries slow, and the cluster enters a spiral merges never win. More hardware buys hours. The real fix is to keep the volatile field out of the index: index the stable fields and hold price in a fast key-value store, applying price filters and sorts at rerank time over the ~1,000 phase-1 candidates. If price must be filterable during retrieval, coarsen it — index a rarely-changing price bucket ("under 5k", "5k–10k"), filter on the bucket, apply exact price at rerank. Both fixes are one insight: a field's churn rate decides whether it belongs in a read-optimized structure.
Requirement flip. The PM says: "Search must be read-your-own-writes. A seller publishes a listing, immediately searches for it, and it must appear. Zero exceptions." Your pipeline has 3–5 seconds of lag. What do you build, and what do you tell the PM?
Don't make the pipeline synchronous — forcing a refresh per write destroys the segment economics and multiplies indexing cost for a rare case. Solve the actual user story in two layers. (1) After publishing, the client holds the new listing ID for ~10 seconds; when the search API sees that ID on the request it fetches that one document from Postgres and splices it into the results if it matches — exact and cheap for the one document the user cares about. (2) The seller's own "my listings" view shouldn't use search at all: query Postgres filtered by seller_id, an indexed lookup that's strongly consistent by construction. Tell the PM: "Global search stays eventually consistent at a few seconds, because making it synchronous costs several times the hardware and slows every query. Your seller sees their listing instantly via two targeted paths; everyone else sees it within five seconds and nobody can tell." Naming the cost of the naive version is the part that reads senior.
What breaks if… you route documents by hash(seller_id) instead of hash(docID), so a seller's listings live together? Name two things that get better and two that get worse.
Better: a query filtered to one seller hits one shard instead of twelve, cutting fan-out cost and tail latency for that pattern; and reindexing one seller's catalog touches one merge pipeline instead of every shard's. Worse: hot shards — the biggest seller may have 50 million listings where the median has 20, so shard sizes and traffic diverge wildly (Part 2.6 in a new costume). And distributed IDF degrades badly: term distributions now differ per shard, so an electronics seller's shard thinks "battery" is common while a shoe seller's shard thinks it's rare, and identical documents score differently depending on where they landed — making the coordinator's merge compare unlike with unlike. That second one produces wrong rankings with no error anywhere, and spotting it unprompted is a strong senior signal.
"Design S3." Two words, and the room goes quiet.
Storing bytes is not hard — a file system stores bytes. What makes this a senior question is the promise printed on the tin: eleven nines of durability, 99.999999999% of objects surviving a given year. Amazon's own gloss makes it feel real: store ten million objects and expect to lose one about every ten thousand years. Underneath sit ordinary disks that die at one to two percent a year — which, at the scale we're about to size, means disks dying every single day. You're being asked to build something that never loses a byte, out of parts that fail on a schedule.
And a second system hides inside the first. Every object needs a row somewhere saying "this key, this version, these checksums, these shards, on these machines." That index is a sharded, replicated, strongly-consistent database with a hundred billion rows in it, and it is every bit as hard as the storage. The famous S3 outage was not a disk problem. It was an index problem.
Two planes, one crux. Let's build it.
Scope out loud first: "I'll build PUT, GET, DELETE and HEAD by key, prefix LIST with pagination, multipart upload, versioning, and light lifecycle rules — tier to cold storage, expire old versions. Out of scope: IAM, cross-region replication, anything query-in-place. Fair?"
Six lines, and the first one already decides the architecture. To find out what it costs, we have to size the thing — which is where this problem stops behaving like every capacity estimate you have ever done.
Here's the move that separates people on this question. Most systems have one scale number. This one has two, and they move independently: the number of objects and the number of bytes.
Bytes → drives. 400 PB logical at the 1.5x redundancy I'll commit to shortly is 600 PB physical; at 20 TB per drive, 30,000 drives. Sanity-check the failures: at a 1.5% annualized failure rate that's roughly 450 dead drives a year — more than one every day. Grow to 100,000 drives and it's four a day, forever. Circle that: drive death is not an incident here, it's a heartbeat. Repair isn't a disaster-recovery feature, it's normal operation.
Objects → metadata rows. 100 billion objects, each needing a row: bucket, key (up to 1 KB of it), version, size, ETag, timestamps, and a manifest listing every chunk with its checksum and location. Call it 1 KB per row — 100 TB of metadata, replicated five ways, and not cold archive: a hot index every request touches. That's a large distributed database before we've stored one byte of anyone's photo.
The punchline: if every customer switched from 4 MB videos to 400 KB thumbnails, the bytes stay flat, the drives are bored — and the metadata plane grows ten times and melts. A genomics customer uploading 5 TB files does the reverse. Two scaling problems, two walls, design against both.
Two more numbers drive decisions. Request rate: ~500K/sec fleet-wide, read-heavy at 5:1 — but fleet-wide is the easy part, it's the per-prefix rate that hurts. Real S3 documents roughly 3,500 write and 5,500 read requests per second per partitioned prefix: one hot prefix is a hot partition (chapter 2.6) with a bucket name on it. Repair bandwidth: rebuilding a dead 20 TB drive reads several times that much from other drives, so repair traffic is budgeted next to customer traffic. Skip it and your first bad day is a network outage stacked on a hardware outage.
PUT /{bucket}/{key} with body + x-amz-checksum-crc32c → 200, an ETag, a version id. GET takes optional ?versionId= and Range:.DELETE /{bucket}/{key} → in a versioned bucket this writes a delete marker; it erases nothing.GET /{bucket}?prefix=photos/2026/&delimiter=/&max-keys=1000&continuation-token=… → up to 1000 keys plus a token.POST …?uploads → upload id; PUT …?partNumber=N&uploadId=… per part; POST …?uploadId=… with the part list to finalize.Two tables in two very different systems. Metadata plane: (bucket, key, version_id) → manifest, holding size, content type, ETag, storage class, and an ordered list of {chunk_id, offset, length, checksum, placement_group}. The primary key is ordered by key, not hashed, and that one choice decides how LIST works — defended in deep dive two. Data plane: no rows at all, just immutable chunks plus a per-drive local index.
One rule pays for itself all chapter: chunks are immutable. Overwriting a key never rewrites bytes in place; it writes new chunks and commits a new manifest. Immutability makes caching safe, checksums meaningful, repair correct, and versioning nearly free.
Write path. The API server authenticates, checks the per-prefix rate limit, and streams the body: cut into chunks, checksum each, erasure-code each chunk into 12 shards, push them in parallel to 12 storage nodes chosen so no more than 4 land in one AZ. Once enough shards are durably on disk, write one row to the metadata shard owning this key — a five-replica Raft group, meaning a majority of the five must agree before anything counts as written. That commit is the moment the object exists — only then does the client see 200.
Read path. GET asks the metadata shard for the manifest — a linearizable read, so it reflects every committed write — then fetches shards in parallel, verifies each checksum, concatenates, streams bytes out. In the common case there is no math at all, for a reason I'll get to. If a shard is missing or slow, pull a parity shard and reconstruct: a degraded read, slower but invisible to the customer.
A library has shelves and a card catalog. The shelves are the data plane: if one collapses you lose a few books. The catalog is the metadata plane: it tells you where every book lives. Burn the catalog and every book is still physically present — and every book is equally lost, because nobody can find anything. That asymmetry is why the planes are separate: shelves need redundancy for the books, the catalog needs redundancy and consistency for the truth. Where the analogy breaks: a real catalog can be rebuilt by walking the shelves. Ours can too, by scanning 30,000 drives, which would take days. Don't rely on it.
Everything else here is assembled from parts you know — a stateless API tier, a sharded database, a queue. The hard part is manufacturing eleven nines out of components that fail daily, at a price a business can pay. Three mechanisms lock together: erasure coding (redundancy at 1.5x instead of 3x), failure-domain-aware placement (so one power event can't take enough shards to matter), and fast, risk-ordered repair (so the vulnerable window is minutes, not days). Miss one and the nines evaporate. The least obvious is the third: repair speed is a durability input, not an operations detail. I'll say that twice; that was once.
The simplest scheme is three copies in three AZs: tolerates two losses, repair is a file copy, reads are trivial — and it costs 3x, which on 400 PB is 1.2 exabytes of hardware, two-thirds of it duplicates. At small scale that's the right answer and you should say so. At 400 PB the finance conversation ends the design.
The alternative, with no field theory. Take 10 numbers and also write down their sum. Erase any one number and you recover it by subtracting the rest: 11 things stored, 1 loss survived, at 1.1x instead of a copy's 2x.
Reed-Solomon generalizes that. Split a chunk into k data shards and compute m parity shards, each a different weighted mix of the data. With k=10, m=4 you store 14 shards and any 10 of the 14 reconstruct the original. Ten unknowns, fourteen equations; pick any ten and solve. Overhead 1.4x, tolerating four losses where 3x replication tolerated two.
Ten friends each memorize one digit of a phone number, and four more each memorize a different arithmetic mix of all ten digits. Any four friends can be unreachable and the remaining ten sit down and solve for the original number, because ten unknowns with ten independent equations has exactly one answer. Where the analogy breaks: real Reed-Solomon works in a finite field, arithmetic that wraps around, so parity stays exactly one byte per byte instead of growing. The intuition — extra equations let you solve for missing unknowns — is exactly right, and it's the sentence to say in an interview.
One property makes your read path fast: Reed-Solomon as deployed is systematic — the k data shards are the original bytes, just cut up. A healthy read is a concatenation with zero decoding; you pay CPU only when a shard is missing.
Now watch a requirement kill the popular answer. I promised to survive losing a whole AZ, and I have 3 AZs: 14 shards over 3 AZs puts at least ceil(14/3) = 5 in one of them. Lose that AZ and 5 shards are gone, but I can only afford 4. Objects become unreadable. Not lost — they return when the AZ does — but "unreadable for four hours" fails the availability promise, and the interviewer will find this hole if you don't.
The general rule is worth memorizing: with A availability zones, at least 1/A of your shards sit in one AZ, so surviving an AZ loss needs parity ≥ n/A, forcing overhead of at least A/(A−1). Three AZs → a 1.5x floor. Four AZs → 1.33x. No code is clever enough to go below it.
| Scheme | Overhead on 400 PB | Tolerates | Survives an AZ (3 AZs)? | Reads to rebuild 1 shard |
|---|---|---|---|---|
| 3x replication | 3.0x — 1.2 EB, ~60K drives | 2 losses | Yes, 1 copy per AZ | 1 (a plain copy) |
| RS 10+4 | 1.4x — 560 PB, ~28K drives | 4 losses | No — 5 shards land in one AZ | 10 |
| RS 8+4 (committed) | 1.5x — 600 PB, ~30K drives | 4 losses | Yes — exactly 4 per AZ | 8 |
| RS 16+4 | 1.25x — 500 PB, ~25K drives | 4 losses | No — 7 shards in one AZ | 16 |
I commit to Reed-Solomon 8+4 across 3 AZs, 4 shards per AZ. Half the cost of replication, tolerates any 4 losses, and an AZ outage leaves exactly 8 shards — exactly enough to serve every read uninterrupted. The cost I state out loud: right after an AZ loss there is zero spare redundancy, so repair must reconstruct into the surviving zones immediately while reads run degraded. I'd rather own that than pretend 10+4 was AZ-safe.
Erasure-coding a 4 KB thumbnail into 12 shards gives twelve 512-byte fragments on twelve machines, each with its own bookkeeping and disk seek — the overhead costs more than the data. Every serious object store fixes this the same way: pack many small objects into large append-only containers, a gigabyte each, and erasure-code the container, not the object. The manifest points at (container_id, offset, length), which is why offsets were in the data model. That gives a two-stage write policy: new writes land as 3 plain replicas in an open container (no encoding on the latency path), and once sealed a background job codes it and drops the replicas. Hot data gets replication's simplicity, settled data gets coding's economics, nobody waits.
Facebook published exactly this pattern at OSDI 2014. Their photo store Haystack sat at an effective replication factor of 3.6x — fine when photos were hot and few, ruinous at exabyte scale, since BLOB access falls off a cliff with age. So they built f4, a separate warm tier using Reed-Solomon 10+4 for 1.4x local overhead plus a cross-data-center XOR, landing at an effective 2.1x. Tier by temperature: replicate the hot stuff for simplicity, erasure-code the cold stuff for money.
You don't derive anything at a whiteboard, but you must show the shape of the argument. For an object to die, 5 of its 12 shards must be permanently gone at once, because any 8 survivors reconstruct it. The chance a specific drive dies today is about 1.5% ÷ 365 ≈ 1 in 25,000, and the chance four more specific drives die inside the repair window is that tiny number multiplied several times over. Multiply four or five small probabilities and you land near eleven nines. That's the whole argument.
Which exposes the lever: each of those terms scales with the repair window, the time between a shard going missing and a fresh one existing. Halve repair time and four terms halve with it, dropping risk by roughly sixteen. Repair speed is a durability input. (Second time, as promised.) Slow repair isn't "behind on maintenance," it's durability loss happening now, silently.
How do you make repair fast? Decluster the placement. If a drive's shard groups always used the same 11 partner drives, its rebuild is capped by 11 machines' read speed — hours or days. Draw each placement group from a randomized set of thousands of candidates instead, and a dead 20 TB drive's rebuild reads come from thousands of peers at once, finishing in minutes. Same hardware, an order of magnitude better durability, purely from placement.
And the caveat strong candidates volunteer: that math assumes independent failures, and real failures are correlated. A shared power unit, a firmware bug in one drive model, a bad manufacturing batch, a botched deploy — these kill many drives together and the multiplication stops applying. Which is why placement is failure-domain-aware rather than merely random, why fleets mix drive models and purchase batches, and why the biggest real-world losses are almost never "too many disks died."
Backblaze does something almost nobody else does: since 2013 they publish Drive Stats, a quarterly report on every drive in their fleet — model, size, days in service, annualized failure rate — with the raw daily data as an open download. It now covers hundreds of thousands of drives and has become the industry's de facto public reference for how drives actually fail, including the finding that failure rates vary enormously between models, not just between vendors. That transparency is the culture point; the architecture point is just as good. Backblaze stores data in Vaults, splitting each file into 20 shards — 17 data and 3 parity — one shard per storage pod across 20 pods, so a file survives three pods failing outright, and they open-sourced their Reed-Solomon implementation in 2015. Two lessons: measure your own fleet's real failure rate instead of trusting a datasheet, and choose failure domains such that the thing that fails together is a thing your code can afford to lose.
A dead drive is the easy failure; it announces itself. The frightening one is a drive that cheerfully returns wrong bytes — bit rot, a firmware bug, a flipped bit in a network card or in RAM. Nothing errors, and the customer finds out months later. So: checksums, created as early as possible and verified everywhere.
Detection has three sources — heartbeats (node gone), the scrubber (silent corruption), and read-time verification failures (a customer stumbled onto it first) — all feeding one queue. Workers always do the same thing: read k healthy shards, verify them, re-derive the missing shard, place it under the failure-domain rules, update the manifest.
Here's the piece that turns an ordinary system into a durable one: the queue is ordered by risk, not arrival time. An object missing 3 shards is one failure from unreadable; an object missing 1 has three failures of slack. FIFO makes the desperate object wait behind a million comfortable ones. So the queue has lanes by shards-missing and the critical lane always drains first. During a big correlated failure — when the queue is longest and durability most at risk — that ordering is worth more than extra hardware.
Time for the honest sentence, which earns points precisely because most candidates skip it: the metadata plane is a sharded, replicated, strongly-consistent, ordered index with a hundred billion rows in it, and it is at least as hard as the storage problem. Google learned this in public — the Google File System kept all file metadata in a single master server's memory, which capped how many files a cluster could hold, and its successor Colossus exists largely to distribute that layer.
Sharding: range, not hash — and I'll pay the cost. Hash-partitioning spreads load perfectly and makes LIST impossible: listing photos/2026/ would have to ask every shard. So I range-partition on (bucket, key), keeping keys sorted, which turns a prefix LIST into a contiguous scan of one or two shards. Each shard is a Raft group of 5 replicas placed 2/2/1 across the AZs, so it survives an AZ loss and still forms a majority, and shards auto-split on size or request rate.
The cost is the hot-partition problem from chapter 2.6: customers love keys like logs/2026-08-14T09:12:03Z, and monotonically increasing keys send every write in the world to the last shard. Auto-splitting on request rate is the answer — split the hot range, hand the halves to different machines, repeat. Real S3 lived this publicly: for years the official advice was to randomize your key prefixes, retired in 2018 once S3 began scaling partitions per prefix automatically.
Strong read-after-write, in one paragraph. The object exists when and only when its metadata row commits — the linearization point, exactly like Dropbox's manifest commit in chapter 3.7. Because the shard is a Raft group, that commit is durable and ordered before the client hears 200, and a leader read under a valid lease reflects every commit that finished before it started. Chunks are immutable, so there are no stale bytes to serve by accident; the only mutable fact in the whole system is "which manifest is current for this key," and it lives in exactly one consensus group. When real S3 shipped strong read-after-write in December 2020 — GETs, LISTs and overwrites, no extra cost, no performance penalty — that was a change in the index layer, not the disks.
LIST, honestly. A paginated ordered scan returning up to 1,000 keys plus a continuation token, costing what the prefix contains rather than what you asked for. Listing a bucket with 10 billion objects is a batch job wearing a request's clothing, and a customer polling LIST in a loop to find new files has built a denial of service against your metadata plane using your own API. Mitigations: throttle LIST harder than GET, support delimiter so scans skip whole "directories," and offer the tools built for the real need — a nightly inventory report delivered as a file (real S3 does this) plus event notifications.
Versioning, deletes and lifecycle fall out of the ordered index cheaply. Version id is part of the primary key, sorted newest-first, so "current version" is the first row in the range. A DELETE writes a delete marker, a tombstone that makes the key look gone while recoverable data sits underneath. Nothing is erased on the request path; reclamation is a background sweep that frees chunks no live manifest references, after a grace period. The deletion principle: deleting late costs money, deleting early costs data — every tie breaks toward late.
Just after 9:37am Pacific, an authorized S3 engineer in us-east-1 was debugging a billing-system slowdown and ran a playbook command to remove a small number of servers. A typo removed a much larger set — and those servers supported two other S3 subsystems. One was the index subsystem, holding metadata and location information for every object in the region, required by every GET, LIST, PUT and DELETE. The other was the placement subsystem, which allocates storage for new objects and itself depends on the index. With too much capacity gone, both needed a full restart. Here's the part to tattoo somewhere: S3 had grown enormously since those subsystems were last fully restarted, the restart took far longer than anyone expected, and safety checks validating metadata integrity ran for hours before the index could serve traffic. Most of the internet's "S3 is down" day was really "the index is starting up." Two lessons: the metadata plane is your availability single point of failure even when your data is perfectly safe, and a recovery path you have never exercised at current scale is not a recovery path, it's a hypothesis. AWS's follow-up split the index into smaller cells and made restart time a tracked property. Bonus detail with a bitter laugh in it: the AWS status dashboard couldn't be updated during the outage, because it ran on S3.
The write path is where you commit to a durability contract, and the contract is one question: what must be true before we say 200?
My ack rule: 200 only after at least 10 of 12 shards are durable on disk, spanning all three AZs; the last shards finish asynchronously as a top-priority repair task. Three parts deserve defending. "Durable on disk" means the drive acknowledged, not that bytes sit in the OS page cache — a rack power event must not turn acked writes into lost writes, so nodes use power-loss-protected devices or an explicit flush. "All three AZs" is what makes the ack survive an AZ dying one second later. And 10, not 8: never acknowledge at exactly k, because there the object has zero redundancy and one later drive failure destroys data the customer was told was safe. Acking at 10 keeps two shards of slack while dropping the two slowest nodes from the latency path — a few milliseconds of p99 against a window of zero redundancy.
Metadata commit last is the discipline from Dropbox's manifest commit (chapter 3.7) and the outbox pattern (chapter 2.1). Shards written but never committed are invisible orphans that GC sweeps up; no customer can observe half an object. A retried PUT after a lost response writes new chunks and commits again — last commit wins, S3's documented semantics, and no interleaving produces a Frankenstein object made of two uploads.
Multipart upload exists because a 5 TB single-stream PUT is a bad bet: one network blip and you restart 5 TB. So: create an upload id, PUT parts in parallel (5 MB minimum each except the last, 5 GB maximum, up to 10,000 parts, with the 5 TB single-object ceiling as its own separate limit on top), each erasure-coded, acked and retryable on its own. Then CompleteMultipartUpload assembles their chunk lists into one manifest and commits it atomically: before that the key doesn't exist, after it the whole object does. The multipart ETag is deliberately not the object's MD5 — it's a hash of the part hashes with a dash and the part count appended, since a real digest would mean re-reading 5 TB. Abandoned uploads get swept by a lifecycle rule, because half-finished parts are invisible bytes customers still pay for.
Two read-path refinements. Since any 8 of 12 suffice, a reader chasing p99 can issue 10 requests and use the first 8 that return — hedged reads, tail-latency insurance for a little bandwidth. And Range reads map to specific chunks via manifest offsets, so a player seeks into the middle of a 4 GB video without reading the first 4 GB.
uploads/ takes 50K writes/sec: per-prefix token buckets, 503 SlowDown with Retry-After, auto-split the range. Throttling delivered as a clear retryable signal is backpressure (chapter 2.7); throttling delivered as timeouts is an outage.What I monitor, and note that the top item is not latency:
Rollout. Storage node software rolls one failure domain at a time, never two AZs at once — a bad binary deployed everywhere at once is exactly the correlated failure erasure coding can't save you from. Encoding formats are read-compatible forever: the manifest carries a format version and new code must read every version ever written, because objects written in 2026 will be read in 2046. And borrow the 2017 lesson — rehearse the metadata restart at current scale, on a schedule.
At 10x. Bytes go to 4 exabytes and drives to 300,000: more of the same, plus a bigger repair-bandwidth budget. The first real wall is the metadata plane at a trillion rows — more Raft groups, more split churn, and a placement service whose own state stops fitting anywhere comfortable. The second is LIST, whose cost grows with objects under a prefix. The third is repair bandwidth, which grows with fleet size while the network doesn't automatically follow. Notice what isn't on that list: the erasure code. Coding parameters scale for free; coordination doesn't.
What the interviewer probes next, roughly in this order:
| Choice | Why | What it costs me |
|---|---|---|
| Reed-Solomon 8+4, 4 shards per AZ | 1.5x instead of 3x; survives any 4 losses and a full AZ | Rebuilding 1 shard reads 8; degraded reads cost CPU and cross-AZ bandwidth; 10+4's cheaper 1.4x is off the table because 3 AZs force a 1.5x floor |
| 3 replicas first, code on seal | No encoding on the write latency path; small objects packed before coding | A background transcode pipeline to operate; recent data is briefly more expensive |
| Range-partitioned metadata | Prefix LIST becomes a contiguous scan of one or two shards | Sequential keys create hot shards; survives only with auto-split on request rate (chapter 2.6) |
| Metadata commit as linearization point | Strong read-after-write falls out; no partially visible objects; earlier steps retryable | The metadata plane becomes the availability single point of failure — see February 2017 |
| Ack at 10 of 12 shards, all 3 AZs | Never acknowledge at zero redundancy; drops the two slowest nodes from the latency path | A brief reduced-redundancy window per write, covered by a top-priority repair task |
| Risk-ordered repair queue | Objects closest to death rebuild first, which is where durability comes from | Starves the routine lane during big events; needs its own backlog alarm |
| Delete = tombstone + async reclaim | Recoverable deletes; no erase work on the request path | Customers pay for bytes after "deleting" them; GC must be fenced against in-flight writes |
Explain to a junior engineer, in five or six sentences, how erasure coding gives better durability than three copies while using half the storage — and why repair speed matters as much as the code you picked.
Take a chunk of data, cut it into 8 equal pieces, and compute 4 extra "parity" pieces, each a different mathematical mix of all 8. Store those 12 pieces on 12 different machines. Any 8 of the 12 rebuild the original — 8 unknowns, and any 8 independent equations solve for them — so losing any 4 costs you nothing. That's 1.5x the storage tolerating 4 failures, where three full copies cost 3x and tolerate only 2. But an object only dies if 5 pieces vanish at the same time, so what really protects you is how fast a missing piece gets replaced: if repair takes ten minutes instead of ten hours, the window for those other failures to pile up is sixty times smaller. That's why storage teams treat repair queue backlog, not disk failure rate, as the number that predicts data loss.
Your company opens a fourth availability zone in the region. Recompute the cheapest erasure code that still survives a full AZ loss, and say whether you'd actually migrate.
With A zones, one zone holds at least n/A shards, so surviving a zone loss needs m ≥ n/A and overhead n/k ≥ A/(A−1). Four zones lowers the floor from 1.5x to 1.33x — say 9+3 (3 per zone; lose one and exactly 9 survive) or 12+4. On 400 PB that's about 67 PB saved, roughly 3,400 drives. Migrate? Not by rewriting existing data: re-encoding an exabyte costs months of repair-class bandwidth, and every byte re-read is a byte that can be corrupted. Switch new containers to the new code, let the manifest's format version handle the mix, and let lifecycle transitions re-encode old data lazily. The senior move is noticing that the migration cost, not the code, is the decision.
10x this: 100 billion objects becomes a trillion. Which plane breaks first, what's the symptom, and what do you do?
The metadata plane, well before the disks. A trillion rows at ~1 KB is a petabyte of hot index across tens of thousands of Raft groups. Symptoms in order: shard-split churn as ranges outgrow their targets; the shard-map service becoming a bottleneck because its own state no longer fits on one machine; leader-election storms during deploys; LIST latency degrading as prefixes span more shards. Fixes: make the shard map hierarchical and cached at the API tier; split cold metadata (old versions, delete markers) away from the hot current-version index; cellularize per bucket (chapter 2.5) so one noisy tenant can't churn a shared shard. Note what didn't break: the erasure code and the storage nodes. Coordination fails to scale; arithmetic doesn't.
Requirement flip: the durability target drops to 99.99% per year — this is a scratch tier for machine-learning intermediates that can be regenerated — but cost per TB must be as low as possible and write throughput must double. What changes?
Almost everything relaxes, and saying so confidently is the point — over-engineering a scratch tier is a real failure mode. Move to a wider, cheaper code (16+2, ~1.13x) or single-AZ placement; dropping AZ survival removes the 1.5x floor entirely. Ack earlier, since a lost object costs a recomputation, not a customer. Scrub quarterly instead of monthly, run repair off-peak, page nobody. Keep two things unchanged: end-to-end checksums, because silently corrupted training data poisons results invisibly, and metadata consistency, because a scratch tier that loses track of what it has is unusable at any price.
"Design YouTube." You almost relax. You know this one: upload to blob storage, write a row, transcode it, stick a CDN in front. Six boxes, twelve minutes — and then thirty minutes of silence you have to fill.
That silence is the interview. Every box on that first diagram is a weekend project. What isn't is the thing living between "user pressed upload" and "someone in Jakarta on a 3 Mbps phone presses play": a machine that turns one 4 GB file into forty files, and a second machine that moves an exabyte a day without the bandwidth bill eating the company. The CRUD is trivial. The pipeline and the delivery are the exam.
So the first thing I say out loud, before drawing anything: which shape of video company are we? YouTube is a user-generated firehose — hundreds of hours of unknown-quality video arriving every minute, almost all of it destined to be watched by nine people. Netflix is a curated catalog — tens of thousands of titles, professionally mastered, each watched millions of times, each known about weeks before anyone can press play. That single difference pushes every later decision in opposite directions: compute spent per title, whether you can pre-position bytes before demand exists, what your cache hit rate looks like. I'll design the UGC version because it's strictly harder, and flag every place where a curated catalog buys a better answer.
Functional, scoped out loud: upload a video, process it, stream it back at adaptive quality, count views. Playback works on a phone, a laptop, and a TV, across a bad network. Live I'll treat as a variant at the end — it shares 70% of the pipeline and breaks the other 30% in interesting ways.
Out of scope, said before the interviewer has to ask: recommendations, search, comments, monetization, copyright matching. Recommendations especially — that's a machine-learning system with its own chapter-length design; here it's a black box behind a 50 ms budget that returns video IDs, and if it times out I show trending. Saying that in one sentence beats twenty vague minutes on embeddings.
Non-functional, committed rather than fished for:
Three calculations, each ending in a decision. Nothing else.
1. Ingest volume. 500 hours/minute = 30,000 hours/hour = 720,000 hours of new video per day. Average upload bitrate across phones and cameras, call it 5 Mbps: one hour ≈ 5 Mbps × 3,600 s = 18 Gbit ≈ 2.25 GB. So ingest ≈ 720,000 × 2.25 GB ≈ 1.6 PB per day of source bytes. Now the ladder multiplies it: five or six renditions per codec, two codecs, total output roughly 3–5x the source. Call it 6–8 PB stored per day, a few exabytes a year, growing forever.
2. Transcode compute. One modern CPU core encodes 1080p H.264 at roughly real time on a sane preset. A full ladder — lower rungs are cheap, the top rung isn't — costs maybe 3 core-hours per source hour for H.264. A modern codec like AV1 costs on the order of 10x that. So the cheap codec alone: 720,000 × 3 = 2.16 million core-hours per day ÷ 24 = ~90,000 cores running flat out, forever, in steady state — before any backlog, before AV1, before a single re-encode campaign.
3. Egress and the bill. Use Netflix for this one, because the numbers are public-ish: 300 million-plus paid memberships × two hours of viewing a day = 600 million viewing-hours/day. At an average 4 Mbps (4K pulls 15, phones pull 1) that's 1.8 GB per hour → ~1 exabyte delivered per day, ~100 Tbps averaged out and several times that at evening peak. Now price it: commercial CDN list prices run several cents per GB, so assume you negotiate brutally down to 1 cent. An exabyte is a billion GB → $10 million a day, ~$4 billion a year, against Netflix's 2024 revenue of roughly $39 billion. A tenth of the company's revenue, spent on someone else's bandwidth.
Now the three decisions those numbers just made, said plainly:
Think of a national newspaper. Writing the story is cheap. The expensive parts are the presses — two million copies won't come off one press overnight, so you split the run across dozens in parallel — and the trucks, which is why you print regionally and drive short distances instead of shipping from one city to every doorstep. Transcoding is the presses. The CDN is the trucks. Where it breaks: a newspaper prints one edition, while we print six editions of every story in two languages of ink, because our readers' eyesight changes minute to minute.
The API's only interesting property is that bytes never travel through my application servers — chapter 1.7's payoff.
POST /videos → {video_id, upload_id, part_urls[]}. Creates the metadata row in state uploading and returns pre-signed URLs pointing straight at blob storage.PUT <presigned part url> — the client uploads 8 MB parts directly to the blob store, in parallel, resumable: a dropped connection on a phone re-uploads one part, not 4 GB.POST /videos/{id}/complete → the blob store assembles the parts; we enqueue the processing job. Idempotent by upload_id (chapter 2.2), so a retried "complete" doesn't launch two pipelines.GET /videos/{id}/playback → {manifest_url, license_url}, where the manifest URL is short-lived and signed so it can't be hotlinked.GET /master.m3u8, GET /v3/seg-00042.m4s. Millions of requests per second that never touch my services at all.Data model, two planes. Bytes plane: object storage holds the mezzanine, the per-rung segments, and the manifests — all immutable objects. Metadata plane: videos(video_id PK, owner_id, title, duration_s, status, visibility, created_at), renditions(video_id, rung_id, codec, height, bitrate, state, path), upload_sessions. Metadata is small — a billion videos × 1 KB ≈ 1 TB — but read on every playback start, so it wants sharding by video_id and a cache in front. I'll commit to sharded MySQL: relational access patterns, a tiny working set, and a sharding problem that is completely solved — which is a story worth telling.
YouTube's video bytes lived in Google's storage systems, but its metadata — the rows describing every video and channel — ran on MySQL, and by 2010 it was buckling. Not from data volume: a billion rows is small. It buckled from connection counts (thousands of app threads each holding a MySQL connection), from single accidental queries that scanned huge tables and took replicas down with them, and from the need to shard horizontally without rewriting every query in the app. So YouTube's engineers built Vitess: a proxy layer presenting one logical database while routing to many MySQL shards underneath, pooling connections, rewriting and de-duplicating queries, and killing expensive ones before they kill a replica. Vitess was open-sourced in 2012, joined the CNCF, and graduated in 2019; it now runs under companies like Slack and Square. The lesson: in a video system the database is not the hard part, and yet the industry's most famous MySQL sharding layer was born there. Mention it in one sentence and move to the pipeline, where the marks are.
Write path: client asks for an upload → gets pre-signed part URLs → pushes bytes straight into blob storage → calls complete → we enqueue a pipeline job → the video sits in processing, visible only to its owner. Before it goes public it passes a gate: malware scan plus content fingerprinting for copyright. That gate is asynchronous and part of the same DAG — make it synchronous and you've put a machine-learning model in the upload path, so uploads now fail whenever the model is slow.
Read path: player asks the Playback API for a manifest URL → the API checks entitlement in the metadata DB, signs a short-lived URL, returns it → from there the player talks only to the CDN, fetching a manifest and then a stream of small segment files. Every one of those is an immutable, cacheable GET. That property is not an accident; it's what makes both deep dives possible.
Both hard parts here are the same problem wearing two hats: bytes at planetary scale — making them, then moving them. Making them is the transcoding DAG: one file becomes dozens of renditions, fast enough to share the link in five minutes, cheap enough that 720,000 hours a day doesn't cost more than the company earns. Moving them is CDN economics: an exabyte a day can't be pulled from a central origin and can't be rented at retail prices. Everything else — API, schema, view counter — is a competent afternoon's work. Spread your 45 minutes evenly across all of it and you'll produce something technically complete that reads as junior. Say which two boxes matter, then live in them.
Start with the question the interviewer is really asking: why transcode at all? Three reasons, each multiplying the output count. Devices: a 2016 smart TV decodes H.264 and nothing else; a 2024 phone decodes AV1 and saves you 30% of the bytes. Networks: the same viewer has 25 Mbps on the sofa and 1.5 Mbps on the train. Screens: pushing 4K to a phone burns money on pixels nobody can see. Multiply codecs × resolutions × bitrates and you get the ladder — the renditions of one video, each a rung the player can climb or fall down.
Now the engineering. A one-hour video costs ~3 core-hours for its H.264 ladder — hours of wall clock on one machine, against a five-minute SLA. So: split, fan out, stitch. Cut the mezzanine into temporal chunks — 30 seconds each, so an hour becomes 120 chunks — and hand every (chunk × rung) pair to a worker. That's 600 independent tasks sharing no state. Across a few hundred workers, 3 core-hours finishes in about 90 seconds of wall clock, plus queueing and stitching. Hours to minutes — and the win grows with video length, which is backwards from most systems and exactly right here: the longest videos need the most parallelism and naturally offer the most.
The subtlety that separates candidates: you cannot cut wherever you like. Compression stores occasional complete frames (keyframes, decodable on their own) and stores everything between them as differences from neighbouring frames. Cut mid-run and the decoder has nothing to start from — the chunk is garbage. So splits must land on keyframes, and to guarantee keyframes exist where you want them, the encoder is forced to emit a fixed keyframe interval in closed GOPs — a GOP is a group of pictures, one keyframe plus the frames that lean on it, and closed means none of those frames reference anything outside the group. One keyframe every 2 or 4 seconds, at identical timestamps across every rung.
That last clause does enormous work. Because every rung has keyframes at the same timestamps, segment #42 of the 360p rung and segment #42 of the 1080p rung cover the same slice of time and each starts decodable. That is what lets a player switch quality mid-video without a stutter: it asks for the next segment from a different rung. Adaptive streaming isn't clever player code, it's a promise made by the encoder. Break the alignment and switches show up as visible glitches or hard stalls — a week-long debugging incident.
Two costs I'd name before being asked. Chunking hurts compression: each chunk restarts rate control and loses lookahead across its boundary, so 5-second chunks produce measurably fatter output than 60-second ones. And the tail task rules the SLA — 599 chunks done in 40 seconds means nothing if one worker is stuck. So: chunks in the tens of seconds, hedged re-execution of stragglers (duplicate any task past p99, take the first result — idempotent writes make that safe), and a quality target per rung rather than per chunk so bitrate doesn't wobble at the seams.
Orchestration is chapter 3.11 collecting rent. The DAG runs on a workflow engine that tracks per-node state, retries with backoff, and dead-letters after a cap. Every task is keyed (video_id, chunk_id, rung_id) and writes a deterministic object path, so a retry overwrites the same object — duplicate execution is harmless, which turns at-least-once into effectively-once with no distributed transaction (chapter 2.2). And crucially, priority lanes. New uploads run in the fast lane; backfill — re-encoding ten years of catalog into a new codec — runs in a starvation-proof slow lane on preemptible capacity, with a weighted scheduler guaranteeing the fast lane a floor of the fleet. Without lanes, one enthusiastic backfill quietly pushes every creator's time-to-playable from 4 minutes to 4 hours, and nobody notices until Twitter does.
Until 2015 Netflix encoded every title with the same fixed bitrate ladder — the same rungs for a cartoon and for a shaky handheld action scene. Their engineers pointed out the obvious in a now-famous tech blog post: content complexity varies enormously, so a fixed ladder over-spends bytes on simple content and under-delivers quality on complex content. Their fix — per-title encode optimization — runs a grid of trial encodes across resolutions and bitrates for each title, scores each with VMAF (the perceptual quality metric Netflix built and open-sourced in 2016, which predicts what human eyes actually rate), and picks the ladder from the best points on the resulting quality curve. Animated and low-motion titles needed dramatically less bitrate for the same measured quality; Netflix reported large average bitrate savings at equal quality, and followed up in 2018 with shot-based optimization that tunes encodes per shot. Here's the senior insight to say out loud: Netflix can burn a hundred times the compute analyzing one title because that title is encoded once and streamed a hundred million times — the analysis is a rounding error per view. YouTube's economics are the mirror image: most uploads are watched a handful of times, so the same math says "encode cheap first, spend big only on the head." Identical reasoning, opposite architecture, purely because the catalog shape differs.
By the late 2010s YouTube's transcoding fleet was one of Google's largest compute consumers, and newer codecs (VP9, then AV1) made the curve worse: better compression saves bandwidth but costs enormous encode compute. Google's answer, described in their 2021 paper on warehouse-scale video acceleration, was Argos — a custom video coding unit, an ASIC built for one job, transcoding, deployed at fleet scale in YouTube's data centers. The reported gain was roughly 20–33x in compute efficiency over their optimized software encoders. The point for an interview isn't chip design; it's that transcode-versus-egress is real money on both sides — you spend compute to make bytes smaller so you spend less shipping them, and past a certain scale the right move is to buy that compute in silicon. So when the interviewer asks "what if transcode cost doubles?", "move the hot path to hardware encoders and pay for software quality only on the popular head" is a senior answer.
Encoding produced pixels; packaging makes them streamable. Each rendition is sliced into segments — self-contained files of 2–10 seconds, 4 s being a good default — plus a manifest (HLS's .m3u8 or DASH's .mpd) listing the rungs and where their segments live. I'll package once as CMAF fragmented-MP4: one set of segment files that both HLS and DASH manifests point at, so I store one copy instead of two. Halving storage and doubling cache hit rate for the price of writing two small text files is the cheapest win in this design.
Segment length is a real trade-off, not trivia. Short segments react to network changes faster and cut live latency, but cost more requests and compress worse (each one starts with an expensive keyframe). Long segments compress better and are gentler on the CDN, but the player is committed to its quality choice for longer. 4 seconds for VOD; 1–2 seconds when latency matters.
ABR — adaptive bitrate — lives entirely in the player, and that's worth defending: all the intelligence sits at the edge so the server can serve dumb static files, which is why one cached copy serves millions of viewers on wildly different networks. The naive algorithm measures throughput (bytes ÷ download time of the last segment) and takes the highest rung that fits. It works, and it oscillates: video traffic is bursty and on-off, so throughput samples lie and the player flip-flops between rungs, which annoys viewers more than steady lower quality would. The refinement Netflix researchers documented at SIGCOMM in 2014, from real player data, uses buffer occupancy as the control signal instead: 30 seconds of buffer means you can afford to step up whatever a noisy sample says; a buffer draining toward zero means step down now. Real players blend both, throughput dominating at startup when the buffer is empty, plus a rule to start low so playback begins fast, then climb.
Adaptive bitrate is a cyclist changing gears on a hill. A beginner watches the road ahead and guesses the slope — that's the throughput estimator, and it's wrong every time the wind shifts. An experienced rider goes by how the pedals feel right now and how much momentum they've banked: plenty of momentum, shift up; legs burning and speed falling, shift down immediately. Buffered seconds of video are that momentum. Where the analogy breaks: a cyclist can shift any time, while a player can only change rungs at a segment boundary — so a 6-second segment is a bike that lets you shift twice a minute.
DRM gets exactly one paragraph. Segments are encrypted once under a common encryption scheme; keys are served separately by a license server speaking each platform's DRM (Widevine, FairPlay, PlayReady) after checking entitlement and device trust. The design-relevant consequence: the bytes stay identical for every viewer, so the CDN caches one encrypted copy and never touches a key. Watermark or personalize each stream instead and your hit rate collapses to zero, taking the entire delivery strategy with it. Security decisions have cache consequences — say that sentence and move on.
An exabyte a day, hundreds of terabits at peak, a $4B/year retail bill. The escape route is one property of the content: popularity follows a brutal power law. A tiny head — a new season drop, the top few thousand videos of the day — serves the overwhelming majority of bytes. Get that head into a cache near the user and almost all traffic is served from a few kilometres away while the origin sees a trickle.
So delivery is a hierarchy: edge cache → regional/shield cache → origin. Target 90–95% at the edge, most of the remainder absorbed regionally, under 1% reaching origin. Two things make that easier here than for a typical web app: segments are immutable, so they cache ~forever with no invalidation logic at all, and the object count per popular title is small.
Put a number on "the head" for a curated catalog and the strategy writes itself. 100,000 hours at ~20 GB per source hour across all rungs and codecs is a 2 PB catalog. If the top 5% drives 80% of viewing — a Zipf-ish assumption I'd validate against real logs rather than believe forever — the hot head is about 100 TB: one dense storage box. Most of what a whole neighbourhood wants tonight fits on a single server you could ship to their ISP. Compare a UGC firehose adding 6–8 PB per day with a head that moves hourly. Same first principles, two different delivery systems.
Then the two companies diverge, and the contrast is the answer to "what's your CDN strategy?"
YouTube can't predict what will be popular — a video that didn't exist an hour ago can be the most-watched thing on earth tonight. So its caching is reactive: Google's global edge network, plus caching nodes placed inside ISP networks, pull content on demand on first request and keep what's hot. The firehose makes pre-positioning nearly useless; popularity is only knowable minutes-to-hours after upload, so the system optimizes for how fast a cache fills, not for what's in it beforehand.
Netflix can predict, and that changes everything — this is the elegant bit of the entire chapter.
You met Open Connect back in the CDN chapter: Netflix builds its own appliances, gives them to internet service providers for free, and fills them overnight with what the neighbourhood will watch tomorrow. Here is the part we skipped, because it only matters once you're the one designing the delivery tier.
First, the storage inside the box is tiered, and the tiering is what makes the economics work. A viewing population's demand is a brutal power law: a small head of titles accounts for most of the bytes, and a long tail accounts for almost none. So an appliance holds the hot head on flash, where random reads are cheap, and the deep tail on spinning disks, where bytes are cheap. One box therefore serves the popular 5% at flash speed and still has the rest locally, instead of needing either an all-flash budget or an all-disk latency penalty.
Second, the fill is a scheduling problem, not a caching problem — and that distinction is the whole senior insight. A normal CDN is reactive: a viewer asks for something, it misses, the edge fetches it from origin during peak hours, at peak prices. Netflix's control plane is predictive: it decides, per appliance, what to load, and it loads it in the middle of the night, when the ISP's network is idle and that bandwidth is effectively free. Peak traffic never generates an upstream fetch at all, because the miss was pre-empted hours earlier. You are converting expensive peak bandwidth into free off-peak bandwidth by knowing your demand in advance.
Third, note the design constraint that makes this legal at all: a curated catalog. Netflix has, at any moment, a countable library and a demand forecast per region — so "what will they want tomorrow" is a solvable prediction. YouTube receives hundreds of hours of new video every minute, most of which will be watched by nobody and a tiny slice of which will explode unpredictably. Same engineers, same money, and the pre-fill strategy is simply unavailable — which is why YouTube's answer is reactive edge caching with aggressive origin shielding instead. When an interviewer asks you to design "video delivery", the first question that separates the two architectures isn't scale. It's whether you can predict what people will ask for.
Two mechanisms hold the hierarchy together when things go wrong. Origin shielding: edges never talk to origin directly, only to a designated shield cache per region, so ten thousand edges missing the same new segment become a handful of origin reads. Request coalescing (chapter 1.6's single-flight, cashed in): inside one cache node, a thousand simultaneous requests for a segment that isn't there yet must produce one upstream fetch, with 999 requests parked until it lands. Skip it and a midnight premiere is a self-inflicted DDoS on your own origin — the classic cache stampede, just with 6 MB objects.
Counting views is a solved problem I refuse to over-solve. The player sends heartbeats; a view counts only past a watch threshold of a few seconds and is deduplicated per session, which quietly kills the cheapest inflation. Events land in a stream; a stream processor keeps an approximate live counter (chapter 3.12's territory — fast, in-memory, occasionally slightly wrong) so creators see movement within seconds, while a nightly batch job recomputes the exact settled count from the event log with proper dedup and fraud filtering and overwrites it. YouTube's famous "301+" freeze was this seam made visible: the fast counter paused while verification caught up. Two counters — one fast, one true — is the right answer here because nobody is paid on the live number, only the settled one.
Recommendations get one sentence and a boundary: a separate ranking service with its own models, called with a 50 ms budget on home-page load, falling back to a cached trending list if it misses. Naming the dependency, budgeting it, and defining the fallback is the whole senior move. The rest is a different interview.
Live reuses the pipeline and deletes its luxuries. The creator's encoder pushes to an ingest endpoint over RTMP (ancient, universal) or SRT (modern, better on lossy networks); the ingest PoP — point of presence, one of your edge sites — transcodes the ladder in real time — no multi-pass, no per-title analysis, no lookahead, one shot at each frame — so you lean on hardware encoders and accept slightly worse compression. The packager emits segments continuously into a rolling manifest and the CDN fans out as before. Keep the last few hours of segments and the DVR window is free: rewinding live is just fetching older segments that are still cached.
| Latency tier | How | Cost | Use it for |
|---|---|---|---|
| 15–30 s (standard) | 6 s segments, 3-segment buffer, plain HLS/DASH | You're half a minute behind Twitter | Concerts, conference streams, anything one-way |
| 6–10 s (tuned) | 2 s segments, 2-segment buffer | More requests, slightly worse compression | Sports where spoilers matter a bit |
| 2–5 s (LL-HLS / LL-DASH) | Partial segments streamed as they encode, blocking manifest requests | CDN must support chunked delivery; rebuffer risk rises | Live sports, auctions, interactive shows |
| < 1 s (WebRTC) | Real-time transport, no segments | Loses HTTP caching — you scale with relay servers, not a CDN, and cost per viewer jumps | Two-way, betting, small-audience interactivity |
The failure mode unique to live is the synchronized start (chapters 2.6 and 2.7 in one place). Five million people press play at kickoff inside ten seconds — and worse, every player then polls the manifest on the same segment cadence, so the herd stays in lockstep forever. Fixes in order of value: cache the manifest at the edge even for one second (a million requests collapse into one origin hit per edge per second); jitter client refresh timers; use LL-HLS blocking playlist requests so the CDN holds the connection and answers when the next part exists instead of being polled; and start every player at a conservative rung with a randomized ramp-up, or a stadium of players all steps up to 4K in the same second and stampedes the same uplink.
Where interviewers push once your diagram is up. A passing senior answer names the mechanism and its cost, not just the component:
| Decision | Why | What it costs me |
|---|---|---|
| Direct-to-blob resumable upload via pre-signed URLs | App tier never carries 4 GB files; mobile retries cost one part, not the whole file | Client complexity, and a second "complete" call that must be idempotent |
| Chunked parallel transcode (30 s chunks) | Hours of wall clock become minutes; hits the 5-minute SLA | Slightly worse compression, keyframe-alignment discipline, straggler chunks decide the SLA |
| Baseline rung now, high rungs lazily by popularity | Compute follows attention; most uploads never justify a full ladder | A video that goes viral in minutes streams at lower quality until the ladder catches up |
| Own edge fleet / ISP-embedded appliances | At an exabyte a day, retail CDN pricing is a tenth of company revenue | Hardware, logistics, ISP relationships — only sane above a very high floor |
| CMAF packaged once for HLS and DASH | One copy of the bytes, one set of cache entries, two manifests | Old devices that only speak legacy transport streams need a fallback pipeline |
| Approximate live view counts + exact nightly settle | Creators see instant feedback; billing and rankings use the true number | Two numbers exist, and they disagree for a while — needs explaining in the UI |
| Standard-latency HLS by default, LL-HLS only where it's worth it | Standard is cheaper, more robust, and rebuffers less | You're ~20 s behind live, which is unacceptable for sports and betting |
Explain to a junior engineer, in five or six sentences, why a video service can't just store the file you uploaded and stream that same file back to everyone.
Your uploaded file is one fixed size and quality, but the people watching are not: one is on a TV with fast fibre, one is on a phone on a train, and one has a device that can't even decode your file's format. So we re-encode the upload into a ladder of versions — several resolutions and bitrates, in more than one codec — and chop each version into a few seconds of video per file. The player fetches those little files one at a time and, between each one, decides whether the network can afford a better version or needs a cheaper one, which is why quality shifts mid-video instead of freezing. That only works because every version has its cut points at exactly the same timestamps, so the pieces are interchangeable. And because the pieces are plain immutable files fetched over HTTP, a cache near the viewer can serve the same piece to a million people, which is the only reason the bandwidth bill is survivable.
Uploads jump 10x overnight, from 500 to 5,000 hours per minute, and stay there. What breaks first, and what are your first three moves in the first hour?
The transcode farm breaks first — everything else is linear and dull. Fast-lane queue age climbs, time-to-playable blows through the 5-minute SLA, and the backlog grows faster than autoscaling can add capacity (and you may hit real quota or hardware limits). Hour one: (1) pause the backfill/re-encode lane entirely to free the whole fleet; (2) degrade the ladder — publish only the 360p/720p baseline so videos become watchable, and queue everything else behind the fresh work; (3) raise the popularity threshold for high rungs and the expensive codec so almost nothing qualifies. All three are shedding quality to protect availability, which is the correct instinct. Then tomorrow: capacity, hardware encoders, and possibly admission control on upload for low-trust accounts. Note what you do not do — you don't drop uploads, because the mezzanine is already safely in blob storage and can be processed late.
The requirement flips: every upload must be available in 4K AV1 within 10 minutes of upload. Redesign what you must, and estimate the damage.
This deletes the single biggest cost-saving decision in the design (lazy popularity-tiered ladders) and replaces it with the most expensive rung on the most expensive codec, for content that mostly nobody will watch. Rough damage: AV1 runs on the order of 10x H.264's encode cost and the 4K rung dwarfs the lower ones, so fleet compute goes up by roughly an order of magnitude — the ~90,000-core steady state becomes something closer to a million cores in software. Real responses: hardware AV1 encoders (this is exactly the Argos-style argument), much finer chunking to hit the 10-minute wall clock (accepting the compression penalty), and pushing back on the requirement with the number — "this costs about 10x our transcode budget to serve a rendition that under 2% of views will ever request; can we make it a creator-tier feature instead?" Bringing a cost estimate to a product argument is the senior move; silently building it is not.
Now add: creators can edit a published video — trim the first 20 seconds, or replace the audio track — without changing its URL or losing its view count.
Two things must be split apart: the identity of the video and the identity of its bytes. Keep video_id stable in metadata (view counts, comments, links all hang off it) and introduce a version_id for the rendition set; playback resolves video_id → current version_id → manifest path. Re-encode into a fresh version directory, then flip the pointer atomically when it's ready — no in-place mutation of segments, because immutability is what makes CDN caching work. Cache handling: the manifest URL must be versioned (or short-TTL) so players pick up the new version; the segment files are new paths, so they can't be stale. Mid-playback viewers keep the old version until they reload, which is fine and honest. And note the trim case can sometimes skip a full re-encode: if the cut lands on a segment boundary, you can republish a manifest that starts at segment N — a manifest edit instead of a compute job, which is the lazy and correct answer when it applies.
What if a requirement flips to sub-second latency for a live auction with 5 million concurrent viewers? Commit to an approach and name what it costs.
Sub-second rules out segment-based delivery, so it rules out HTTP caching, so it rules out the CDN economics this entire design rests on. I'd commit to a hybrid: WebRTC (or a low-latency proprietary transport) for the interactive bidding surface, delivered through a tree of relay servers, while the video runs on LL-HLS at 2–3 s for everyone who is only watching. Costs and honesty: relay fan-out scales with viewers instead of being free at the edge, so cost per viewer is roughly linear rather than flat — 5M concurrent is a serious bill and a serious ops problem; a lost relay drops sessions rather than being retried away; and you now run two delivery stacks. I'd also push on the requirement, because in most auctions what actually needs sub-second delivery is the bid state (a few hundred bytes) and not the video — sending bid updates over a persistent connection while video stays 2 s behind is dramatically cheaper and usually indistinguishable to users. Separating the small urgent thing from the big heavy thing is the general form of this answer.
"Design a stock exchange."
Your hands already know what to do. Load balancer. Shard by symbol. Kafka between the tiers. Cassandra for the orders, Redis in front, three regions for availability. You've drawn this machine twenty times in this book and it has been right nearly every time.
It is wrong here — not slightly wrong, wrong in the direction it points. Every one of those boxes exists to spread work across more machines, and spreading work costs you a network hop, a lock, or a consensus round. In an exchange, matching one order takes about two hundred nanoseconds of actual arithmetic. Asking a second machine its opinion about that order costs ten thousand nanoseconds. The coordination costs fifty times more than the work being coordinated.
So here's the thesis, and you should say it out loud in the first three minutes, because it is most of the answer: an exchange is the one system where you go faster by using fewer machines. The hot path is a single thread on a single box, and the entire architecture around it exists to keep that thread fed, deterministic, and recoverable. Everything you learned about horizontal scale still applies — just not to the part that matters.
Scope out loud: "I'll build the continuous trading session for one equities venue — limit and market orders, cancel and modify, price-time priority matching, a real-time market data feed, in-path pre-trade risk, and a regulator-grade audit trail. Out of scope: the opening and closing auctions, which are a different batch algorithm, plus clearing and settlement. Fair?"
Functional: limit orders (buy 500 AAPL at 100.01 or better) and market orders; cancel and modify; price-time priority matching — best price wins, and at the same price whoever arrived first wins; market data to thousands of subscribers; execution reports back to the owner; pre-trade risk checks; and audit good enough to reconstruct any moment of any day on demand.
Non-functional, with numbers committed — these build the machine:
Notice what nobody asked for: "scales to N machines." Adding it uninvited costs us two requirements that were asked for.
This estimation section doesn't size a fleet. It decides whether a fleet should exist.
Compute per message. A million messages a second is one message every 1,000 nanoseconds. A core at ~3 GHz gives about 3,000 clock cycles in that window. What does adding a resting limit order cost? An array index (a subtract and a shift), a node appended to a linked list, an aggregate quantity bumped, a hash entry written so a future cancel can find it, one event dropped in an outbound ring buffer. A few dozen writes, nearly all to memory already in cache: roughly 100–300 ns. One core at peak volume is about 20–30% busy.
Say the conclusion out loud, because it is the design: one core, doing integer work on cache-resident data, absorbs the venue's entire peak with room to spare. Now price the alternative.
| Operation | Rough cost | Multiple of one match |
|---|---|---|
| Match an order on a warm single thread | ~200 ns | 1x |
| Uncontended lock acquire + release | ~20 ns | 0.1x |
| Cache line ping-ponging between two cores | ~100–200 ns | ~1x per contended write |
| Network round trip, kernel bypass, same rack | ~5–10 µs | ~50x |
| Network round trip through the kernel TCP stack | ~30–50 µs | ~200x |
| One consensus round (Raft-style, 3 nodes) | hundreds of µs | ~1,000x |
| fsync to NVMe | ~100 µs to several ms | 500x–10,000x |
Two rows need a word of explanation. Kernel bypass means the network card hands packets straight to your process's memory, skipping the operating system's networking code entirely — that's the difference between the 5–10 µs row and the 30–50 µs one. And a consensus round is what it costs to make several machines agree on one fact before anyone acts on it.
Read the right column once more. Distribution is always a trade: you buy throughput by paying coordination. Here the work is so cheap that the coordination costs more than the work it distributes. So don't buy it.
Memory. An order record is ~64 bytes — id, account, symbol, integer price, quantity, sequence number, two list pointers. Twenty million resting orders is 1.3 GB, and ten thousand books at a few thousand price slots each is under another gigabyte. The whole market fits in RAM on one ordinary server, and the few hundred symbols carrying most of the volume stay resident in L3 cache. Conclusion: no database in the hot path. Not a fast one, not an in-memory one behind a socket. State is memory, in-process, on that core.
Durability volume. 200K msg/sec over a 6.5-hour session is ~4.7 billion messages/day; at 64 bytes that's ~300 GB/day of append-only journal — 13 MB/s average, 64 MB/s at peak. One NVMe drive laughs at that. So durability isn't a throughput problem here. It's a latency problem: writing the bytes is free, waiting for the disk to confirm them is not. That decides deep dive three.
A busy deli with one counter. Customers take a numbered ticket at the door and one very fast server works the tickets in order. Someone says: add three more servers, it'll be three times faster. But each sandwich takes eight seconds and the servers would spend their time shouting across the room — "did you take the last pastrami? whose turn is it?" — and the shouting takes longer than the sandwiches. Three servers is slower, and worse, customers start getting served out of turn. The fix is a ticket machine nobody can bypass and one server who never stops moving. Where the analogy breaks, usefully: the deli can put a second person on the drinks counter, because drinks and sandwiches never interact. That is exactly how a real exchange scales — split by symbol, never split a symbol.
Two protocols, both binary, because parsing text costs microseconds we don't have. FIX is the industry lingua franca and we accept it at the edge for compatibility, but the latency-sensitive path speaks fixed-width binary — every field at a known byte offset, no allocation to decode. Nasdaq's OUCH (order entry) and ITCH (market data) are the canonical examples of this shape.
NewOrder{client_order_id, account, symbol, side, qty, price_ticks, order_type, tif}, Cancel{client_order_id}, Replace{client_order_id, new_qty, new_price_ticks}.Accepted, Rejected{reason}, Fill{qty, price, exec_id}, Cancelled — each carrying the sequence number that caused it, which is the client's handle for reconciliation.AddOrder, OrderExecuted, OrderCancelled, Trade — each with a feed sequence number, plus a snapshot channel for late joiners and gap recovery.The data model is where candidates reach for tables and lose the thread. There are none. The data model is a memory layout. An Order is a fixed 64-byte struct from a pre-allocated pool. A PriceLevel is an aggregate quantity plus head and tail pointers into a FIFO list. A Book is a flat array of price levels indexed by tick, one per symbol. Persistence is one append-only log of inbound messages.
That log is about to do far more work than "persistence" suggests. Hold onto it.
Order path. A client's packet lands on a gateway, which authenticates the session, normalizes the message into the internal fixed-width format, stamps a hardware arrival timestamp from the network card, and hands it to the risk gate. Risk does a handful of in-memory comparisons and either rejects it right there — the client gets a reject in microseconds and the order never existed as far as the market is concerned — or passes it on. The sequencer stamps a global sequence number and publishes. The engine consumes that stream in order, mutates the book, emits events. Publishers turn those events into market data and execution reports.
Read path. There is no query path; nobody runs SELECT * FROM order_book. The read path is the feed: a continuous multicast of every change plus a snapshot channel to bootstrap from, and clients maintain their own copy of the book by applying it. Name that decision — we push state changes instead of serving queries, because ten thousand subscribers polling the top of book would be a self-inflicted denial of service, and because everyone must see each change at the same instant.
Now claim the agenda instead of waiting for it: "Three things are genuinely hard — the sequencer and why one authority is non-negotiable, the matching engine's mechanical design, and how you fail over a single-threaded machine without changing history. Sequencer first, because everything else depends on it."
Price-time priority makes "who was first" a legal fact, not an implementation detail. If two machines can ever disagree about which order arrived first, you have two versions of the truth and no way to arbitrate — and the loser sues. So the design must contain exactly one component whose job is to decide order, and every other component must be a deterministic replay of that decision. Get this right and matching, recovery, standby, audit and testing all fall out of it for free. Get it wrong and no amount of clever matching logic saves you: you've built a system whose central promise is unverifiable.
The sequencer is almost embarrassingly simple, and that's the point. It receives messages from every gateway on one input path, assigns each a strictly increasing 64-bit sequence number and a timestamp, and immediately publishes. No business logic — it doesn't know what a limit order is. It allocates nothing, branches almost never, and can be a few hundred lines, or at the extreme an FPGA stamping packets in the network card. "Simple enough to be provably correct" is the design goal, because this is the venue's single point of truth.
That stamped output — the sequenced stream — is the system of record. Not the engine's memory. Not any database. The log. The engine applies it and produces trades; the warm standby applies the identical stream and holds identical state; the journal persists it so state can be rebuilt from zero; surveillance reads it minutes later hunting manipulation — spoofing (posting orders you never intend to trade, to fake demand) and wash trades (buying from yourself to fake volume); and the regulator gets it, because the sequenced log is the audit trail — not a report generated from a database, but the literal inputs, in order, replayable to reproduce any millisecond of the day. This is event sourcing — storing the stream of inputs rather than the current state — pushed all the way to its end, and it is the replay superpower of chapter 2.9 taken literally: the log is the database, and everything that looks like state is a cached projection of it.
"But that's a single point of failure." Yes — deliberately, and answered twice: by making it small enough to almost never fail for its own reasons, and by a hot-warm pair with epoch fencing. What you must not do is dodge that by having two sequencers agree via consensus per message — a Raft round trip per order, a thousand times our budget. The single writer isn't a compromise we tolerate. It's the design.
Around 2010, LMAX was building a retail trading venue in London and hit exactly this wall. Their first attempts used the standard playbook: stages connected by concurrent queues, work spread across threads. The queues were the problem — every hand-off meant contended locks, cache lines bouncing between cores, and unpredictable pauses. Martin Thompson, Mike Barker and Dave Farley threw the model out and built the Disruptor: a pre-allocated ring buffer of fixed-size slots, coordinated by sequence counters and compare-and-swap, with no locks and no garbage created on the hot path. Cache-line padding kept unrelated counters off the same 64-byte line so cores stopped fighting over them. The business logic itself ran on a single thread, all state in memory, made durable by journaling and replicating the input events. Their published result was over 6 million orders per second on one thread in the JVM — a language people insisted was too slow for this — beating the multi-threaded design it replaced. LMAX open-sourced the Disruptor and it won a Duke's Choice Award in 2011. The lesson generalized far past finance: for small units of work, contention costs more than the work.
Two things to get right: the thread, and the data structure.
The engine loops: read the next sequenced message, apply it, emit events, repeat. Because it's alone, there are no locks, no atomics, and no memory barriers on the hot path — the biggest cost in most concurrent code simply doesn't exist.
To hold single-digit microseconds, that thread needs help from the machine. This is "mechanical sympathy": writing code that matches how the hardware behaves. Allocate nothing at runtime — every order object, event slot and buffer comes from a pool built at startup, because in a garbage-collected language one 10-millisecond collection pause is a thousand times your entire latency budget. That's why matching engines are written in C++ with pre-allocated arenas, or in a very particular dialect of Java that creates no garbage at all: flat primitive arrays, object pools, off-heap buffers. LMAX proved the Java path works. Keep hot data contiguous so the prefetcher guesses right. Pin the thread to a physical core and tell the OS scheduler to leave that core alone. Busy-spin rather than block, because parking and waking a thread costs microseconds and spinning costs only electricity. In this one corner of computing, burning a core at 100% doing nothing is correct.
Prices aren't continuous — they come in ticks, one cent for a US equity above a dollar. So a price is really an integer index, and the book is a flat array of price levels: index = (price - base) / tick_size. Each slot holds an aggregate quantity plus the head and tail pointers of a FIFO queue of orders resting at that price.
Why an array rather than the balanced tree most textbooks reach for? Contiguous memory is cache-friendly and pointer-chasing is not; reaching a known price is arithmetic, not a search; best bid and best ask are cursors that step outward when a level empties; cancel is O(1) via a hash map from order id to node, then an unlink. Say the cost before the interviewer finds it — memory scales with the price range, not the order count, so a wide-range instrument wastes slots. Two things rescue it: exchanges already impose price bands, so an order 20% from the reference is rejected anyway, and unusual instruments can fall back to a sparse map at a small latency cost.
Walk one by hand. The ask side is the diagram above: 500 at 100.01 (order A 300, then order B 200), 1000 at 100.02, 2000 at 100.03. In comes buy 2000, limit 100.03.
The engine walks outward from the best ask: take 300 from A, 200 from B, level 100.01 empties so advance the cursor; take 1000 from C at 100.02, advance again; take 500 from D at 100.03, leaving D with 1500 resting. Four fills, eight execution reports, four trade prints, a new best ask. Had the incoming order been 5000, the leftover 1000 would rest at 100.03 behind D — new arrivals always join the back.
Two details that mark someone who has thought about this. Time is the sequence number, not the wall clock: clocks drift, sequence numbers don't, and every comparison in the engine uses the sequence. And modify semantics: reducing quantity keeps your place in the queue, but raising quantity or changing price sends you to the back. That's not arbitrary — it stops people holding a good queue position with a token 100 shares and inflating it the instant it's about to trade. A market order is the same sweep with no price limit, plus a protective band so a thin book can't fill someone at an absurd price.
Single-threaded, all in memory, no database. So what happens when that box dies mid-session? The answer rests entirely on the property we've been building toward: engine state is a pure function of the sequenced stream. If that holds, any machine that has consumed the stream has the same book, and recovery isn't a restore — it's a replay.
Determinism isn't free. It's a discipline the whole team holds. No wall clock in business logic — time comes from the timestamp inside the sequenced message, never a local clock call, or an expiry check fires at a different point on replay and you've rewritten history. No floating point for prices; integer ticks everywhere. No iteration over anything unordered — a hash map's iteration order is a bug waiting for a rehash. No randomness, no thread scheduling, no I/O results influencing outcomes; if something outside must influence a decision, it enters as a sequenced message like everything else.
And it's tested as a property, not assumed: replay yesterday's production journal through the build and byte-compare the final book state and every output message against what production actually emitted. Any difference is a defect until proven to be an intended rule change. Saying that out loud lands hard in an interview, because it shows you know determinism decays silently unless something checks it.
fsync costs 100 µs to several ms — 500 to 10,000 times the cost of a match — so waiting for the disk on every message isn't available to us. Instead: the sequencer publishes a message only once it is in memory on at least two of three journal nodes in different racks, about 10–20 µs on a kernel-bypass network. Those nodes write to disk continuously, but asynchronously, behind the acknowledgment.
The exposure, stated honestly: if all three journal nodes lose power in the same instant, the last few milliseconds of sequenced-but-unwritten messages are gone. I accept that and cover it — separate racks and power feeds so correlated failure needs a data-centre-level event, plus end-of-day reconciliation against clearing, which the industry already runs. What I won't do is pay 100 µs on every order to prevent a failure mode that requires a building to go dark.
Well under a second, no lost orders, no invented ones — and this only works because of determinism. If the standby could produce even slightly different output, promoting it would silently rewrite the market's history.
Facebook's IPO was the biggest in Nasdaq's history, and it broke on the one component that isn't the continuous matching engine: the opening cross, the batch algorithm that computes a single opening price from all accumulated interest. The cross calculates a price, then checks whether any orders arrived or were cancelled during the calculation — and if so, recalculates. With that much volume, cancellations kept streaming in faster than the loop could finish, so it recalculated and never converged. The open was delayed roughly half an hour, and about 30,000 orders sat in limbo for more than two hours while traders couldn't learn whether they owned Facebook stock. Nasdaq later paid a $10 million SEC penalty — the largest ever levied against an exchange at the time — and set up a fund of about $62 million to compensate member firms. The engineering lesson is this chapter's thesis seen from the other side: a race between "compute the result" and "the inputs keep changing" is exactly what one ordered stream and a deterministic consumer eliminate. Take a sequence-numbered snapshot of the inputs, compute from it, and later messages belong to the next round instead of poisoning the current one.
One trade becomes thousands of outbound messages. At 200K events/sec with 500 subscribers that's 100 million messages/sec of egress — a genuine horizontal-scale problem, and emphatically not the engine's job. So the engine writes each event into a local ring buffer and moves on, never blocking, and a separate fleet of feed handlers consumes that buffer and does all per-client work: entitlements, TCP sessions, throttling, geographic relay, and conflation — collapsing several updates for the same price level into just the latest one, so a slow client gets a correct current picture instead of a correct but hopeless backlog.
Chapter 2.7 taught backpressure as "slow the producer." Here that inverts. The producer is the market; it cannot be slowed, and a design where one subscriber's congested TCP window can stall matching is a design where one slow customer halts a national market. The rule becomes shed subscribers, never shed orders: a consumer that can't keep up gets conflated updates, then gets disconnected and told to re-bootstrap from the snapshot channel.
The feed is UDP multicast — one packet reaching a thousand listeners is the only way to serve thousands of subscribers at the same instant — so it can drop packets. Handled like everything else here: every feed message carries a sequence number, clients detect their own gaps, and a separate retransmission service serves the missing packets from the journal. The engine is never asked to repeat itself.
Market data is a stadium announcer, not a customer service desk. The announcer says the score once over the PA and eighty thousand people hear it at the same instant. If someone in row 40 was buying a hot dog and missed it, the announcer doesn't stop the game to repeat — that's what the scoreboard (the snapshot channel) and the info booth (the retransmit service) are for. The moment one confused spectator can interrupt the announcer, everyone else's game stops. Where it breaks down: a stadium doesn't much care if row 40 stays confused, whereas an exchange has a duty to make recovery genuinely available — hence a real snapshot channel, not a shrug.
Risk checks must happen before an order can trade — not an architectural preference but regulation; in the US the SEC's Market Access Rule requires brokers to have pre-trade controls. And every check sits directly in the latency path participants measure you on. That tension is the interesting part.
Resolve it by keeping every in-path check O(1) and local, in gateway memory: price inside the allowed band around the reference (the fat-finger guard — a buy 40% above the last trade is a typo far more often than a strategy), size under the account's cap, notional under its remaining headroom. A few integer comparisons, a couple of microseconds.
The trap is credit. An account's headroom is shared state across every gateway it can reach — and shared state means coordination, which we spent the chapter banning. I'll commit to a combination: route all of one account's orders to one gateway so the state has a single owner (partitioning, not coordination), and for large accounts genuinely needing several gateways, lease each gateway a slice of the credit — the leased-ID-block trick from chapter 3.1 — reconciled asynchronously by a slower risk service. The cost is honest: leases over-reserve, so an account can occasionally be rejected while it technically had room elsewhere. I take that side, because the other error is uncontrolled exposure, and that one has bankrupted firms.
Everything that tolerates delay lives off the path. Surveillance for spoofing, layering and wash trading is just another consumer of the sequenced stream running minutes behind — same data, zero latency impact. State the principle: anything that can be seconds late must be seconds late.
The sequencer decides order, so fairness reduces to a physical question: does everyone have an equal shot at reaching it first? That's why exchanges cut every colocation cross-connect to the same physical length — the cabinet nearest the engine gets the same coil of extra fiber as the one across the hall, so nobody's photons take a shorter trip. Fairness measured in meters. Beyond that: one pipeline for everyone, no premium express lane, no order type that skips a stage.
IEX — the venue at the centre of Michael Lewis's Flash Boys — took the opposite approach to latency: it deliberately added some. Between its point of presence and its matching engine sits a physical coil of roughly 38 miles of fiber optic cable, wound into a box, delaying every message by 350 microseconds in each direction. The reason is a specific unfairness. When a large order sweeps several venues it reaches the nearest one first, and a fast trader who sees that fill can race ahead to the other venues and move prices before the rest of the order lands. IEX's delay means that by the time an order emerges from the coil, IEX's own pricing logic has already processed the public market data and repriced its pegged orders — orders that don't name a fixed price but track the wider market's best bid or offer automatically. So the racer arrives to find the price already moved. It was fiercely contested; opponents argued a deliberate delay violated the requirement that quotations be "immediately" accessible. The SEC approved IEX as a national securities exchange in June 2016, taking the view that a delay under a millisecond is too small to matter. The point for your interview: latency and fairness are different goals, and a venue sometimes chooses fairness. Know which one your design optimizes, and be able to say why.
You cannot canary a matching engine: there is exactly one, and two versions side by side would produce two truths. So the release process differs in kind. First, replay yesterday's full production journal through the new build and byte-compare every output message against what production emitted; any diff is a stop-ship unless it's an intended behaviour change, in which case only the expected messages may differ. Second, ship the build to the warm standby and let it run a full session as a shadow — same input stream, output compared and discarded. A day of production traffic with zero blast radius. Third, promote in a planned window, with the old build sitting warm as the rollback.
Ten million messages a second is 100 ns per message — below the cost of the work itself. So, honestly: the single thread breaks first, somewhere in the low millions of messages per second for realistic book logic. Say that plainly rather than defending the design.
Then the escape hatch, which doesn't contradict the thesis: partition by symbol. Run N independent pipelines — sequencer, engine, standby, publishers — each owning a disjoint set of symbols. AAPL and MSFT never interact, so the shards never talk. That is distribution without coordination, the real rule this chapter teaches: not "never distribute" but "never coordinate." Name the price — cross-symbol atomicity disappears, which the first exercise below picks apart. Second to break, and earlier in wall-clock terms: market data fan-out, which scales the ordinary way, with more feed handlers and relays.
A passing answer says "single-threaded" in the first five minutes and defends it with numbers rather than vibes; names the sequencer as the crux; explains recovery as replay, not restore; and volunteers the failover walk unprompted. What sinks candidates: reaching for Kafka and Cassandra on the order path, or hand-waving "we'd use consensus for ordering" without pricing the round trip. Expect these probes:
| Choice | Why | What it costs |
|---|---|---|
| Single-threaded matching core | Matching is ~200 ns; any coordination costs more than the work | Hard ceiling in the low millions of msg/sec; one slow code path stalls everything behind it |
| One sequencer as ordering authority | "Who was first" must have exactly one answer, and consensus per message is 1,000x too slow | A single point of failure that must be fenced and failed over correctly, plus one extra hop per message |
| State in RAM, no database in path | The whole market is ~1.3 GB; a hop to any store is 50x the match cost | Recovery depends entirely on replay; a determinism bug becomes a production correctness bug |
| Durable via replicated memory, fsync async | fsync is 500–10,000x the match cost | Simultaneous loss of three racks loses the tail; covered by reconciliation, not prevented |
| Price-level array indexed by tick | O(1) price access, cache-friendly, cancel is an unlink | Memory scales with price range, not order count; wide-range instruments need a sparse fallback |
| Scale only by symbol partition | Disjoint symbols never interact — distribution without coordination | No atomic cross-symbol orders; multi-leg trading needs its own instrument and book |
A junior asks: "Everything we build scales by adding servers. Why is the exchange's matching engine one thread on one machine? Isn't that a bottleneck waiting to happen?" Answer in five or six sentences.
Adding servers is a trade — you get parallel work and you pay for it in coordination: locks, network hops, agreement rounds. Matching an order is tiny, about two hundred nanoseconds of integer arithmetic on data already in cache, while asking another machine anything costs five to ten microseconds. So the coordination would cost fifty times more than the work it coordinates, and we'd end up slower. On top of that, an exchange must guarantee orders are handled in exactly the order they arrived, and the only cheap way to have one true answer to "who was first" is to let one component decide it and one thread apply it. So we spend one core, keep the whole market in memory, and put all our engineering into keeping that thread fed and perfectly replayable. When we genuinely need more capacity we split by symbol into independent pipelines — distribution without coordination, which is the real rule: never coordinate, rather than never distribute.
Now add spread orders: a client wants to buy 100 of symbol X and sell 100 of symbol Y atomically — both legs or neither. Your symbols are partitioned across independent engines. What breaks, and what do you commit to?
What breaks is the property that made partitioning free: the shards never talked. An atomic two-leg fill spans two engines, so you're back in distributed-transaction territory (chapter 2.1) — a coordinator, a prepare phase, orders pinned on both books while you wait, and a hard question about whether a pinned order is still tradable. That's tens or hundreds of microseconds of coordination in the one path you spent the whole design keeping local. The industry's actual answer, and the one to commit to: make the spread its own instrument with its own order book on a single engine, traded as "X minus Y" with normal price-time priority, and run a separate implied-pricing process linking it to the outright books. Everything atomic stays inside one thread. The cost is fragmented liquidity plus the complexity of implied orders.
Peak volume goes 10x, to 10 million messages/sec. Walk the failure order: what breaks first, second, third — and what do you do about each?
First the matching thread: 10M/sec is 100 ns per message, below the cost of the work, so the input ring buffer fills and queue depth — not latency — shows it first. Fix: partition by symbol into pipelines sized to stay under roughly 2M msg/sec each. Note the honest limit while you're there — the hottest single symbol sets the floor, and if one symbol alone exceeds a thread, no partitioning saves you. Second: market data egress, 10x events on the same subscriber count — add feed handlers and relays, offer conflated feeds. Third the journal: 3 TB/day of sequential writes is still fine, but cold-recovery replay time becomes unacceptable, so snapshot cadence must rise — every few million sequence numbers, taken by a separate process from a copy so the engine never blocks. Fourth, quietly: credit leases exhaust faster and false rejects climb, so lease sizes need retuning. The shape of the senior answer is the ordered list plus the metric that reveals each item.
Requirement flips: a regulator rules that orders must be processed in the order they arrived at the gateways, by hardware timestamp, rather than the order they reached the sequencer. What changes, and what does it cost?
You've replaced a decision with an agreement, and that's expensive. The sequencer can no longer stamp on arrival; it must buffer for a window long enough that any earlier-timestamped message has certainly arrived, sort the window, then release in timestamp order. That adds a mandatory delay of at least the worst-case gateway-to-sequencer spread — tens of microseconds — plus a policy for stragglers arriving after their window closed (drop them, or accept out of order and break the promise). It also makes clock synchronization a correctness requirement rather than a compliance one: every gateway must agree to well inside the window, drift becomes a market-integrity bug, and a slightly fast clock is now worth money to its owner. Honest summary: you trade a cheap, unambiguous, physically-enforced ordering for an expensive one that depends on trusting many clocks. If forced I'd build the buffer-and-sort window, publish its size, and monitor straggler rate as a first-class metric — but I'd argue hard for the original design with equalized cable lengths as the fairness mechanism.
Your nightly determinism test fails: replaying yesterday's journal produces a book differing from production in one symbol, by one order that should have expired and didn't. Name three plausible causes and say which you'd check first.
(1) Business logic read a wall clock — an expiry check calling the machine's clock instead of using the timestamp in the sequenced message, so "now" differs on replay. (2) Something entered the engine outside the sequenced stream — an operator action or a symbol-config change applied by a side channel, so replay never sees it. (3) Iteration over an unordered collection — a hash map of pending expiries rehashed differently, changing which order won a tie. Check (1) first: it's the most common by a wide margin, it's greppable in minutes, and it's the one that looks harmless in code review. The deeper point is that all three are the same bug — a hidden input the log doesn't contain. Determinism is the property that the log is the complete input, and every violation is something sneaking in from the side.
A user types "explain quantum tunnelling like I'm twelve" and hits enter. Somewhere in a datacenter, eight GPUs that cost more than a house spend twenty seconds producing four hundred words — and they are doing the same thing for a hundred and seventy other people at once. Give this user more of the machine and everyone else gets less. There is no cache to hide behind, no CDN edge, no read replica. Every token is computed fresh, on hardware you rent by the second.
Now the interview version: "Design the serving infrastructure for a large language model." Most candidates reach for shapes they know — load balancer, autoscaling group, Redis, a queue — and build something that would be fine for a REST API and is quietly nonsense here. They autoscale on CPU. They put a response cache in front of "the model." They quote a p99 for an answer that arrives one word at a time over thirty seconds.
The interviewer wants one thing: do you understand that a GPU serving an LLM is a memory allocator with a scheduler bolted on, and that how many users you serve, how fast their words appear, and what your monthly bill is all fall out of one resource you cannot buy your way around. Let's find it fast.
Functional, committed in two minutes:
Out of scope, said out loud: training, fine-tuning, model quality, agents, billing invoices.
Non-functional, with numbers I'll defend all chapter:
Four terms. Use them precisely and the interviewer relaxes; say "latency" unqualified and they stop trusting the rest.
Hold two facts together: throughput and per-user latency are in direct tension, and the dial between them is batch size. That tension is the chapter.
A request is not one workload. It is two, back to back, with opposite bottlenecks. I'd draw this before drawing a single box.
Prefill. The model reads your whole prompt in one forward pass — all 1,500 tokens together, as big matrix multiplications. GPUs love this: dense math, arithmetic units saturated. Prefill is compute-bound. Its output is the first token plus the KV cache: saved attention state for every prompt token, kept so later tokens don't re-read the prompt from scratch.
Decode. Now it generates, one token at a time, strictly sequential — token 12 needs token 11 as input. Each step does tiny arithmetic, but to do it the GPU must read all the weights out of memory. For a 70B model at 16-bit that's 140 GB of memory traffic per word fragment. Decode is memory-bandwidth-bound.
One number makes decode's problem vivid. Count arithmetic intensity — how much math you do per byte you drag out of memory. At batch size 1, decode does two operations per parameter and reads two bytes per parameter, so about 1 FLOP per byte. A modern datacenter GPU can do roughly 300 FLOPs in the time it reads one byte. So a single-user decode step uses well under 1% of the chip's arithmetic — a Formula 1 engine idling in traffic.
The fix writes itself: read the weights once, use them for many users at the same time. Batch 64 requests and that same 140 GB read produces 64 tokens instead of 1 — and intensity rises to roughly 1 FLOP per byte per sequence. Throughput grows almost linearly with batch size until that intensity catches up with the hardware's 300, somewhere around batch 200–300 here. Batching isn't an optimisation. It is the business model.
Prefill is a chef reading the whole order ticket for a table of eight in one glance. Decode is the chef walking to the walk-in freezer, opening it, and coming back with one pea. Then walking back for the next pea. The walk is memory bandwidth; the pea is the token. This kitchen only makes money if the chef carries peas for twenty tables per trip — that's batching. Where the analogy breaks: our chef must decide the whole trip's cargo before starting, and the freezer has a hard limit on how much can be carried at once, which is exactly the constraint we're about to meet.
Two calculations. One sizes the fleet; the other changes the architecture, so I do it first and out loud.
A node of 8 GPUs at 80 GB each: 640 GB. Where does it go?
How much KV does one conversation eat? Memorise the formula; interviewers ask you to derive it.
KV bytes/token = 2 (K and V) × layers × kv_heads × head_dim × bytes_per_value
70B-class model with grouped-query attention:
= 2 × 80 layers × 8 kv_heads × 128 head_dim × 2 bytes
= 327,680 bytes ≈ 320 KB per token
320 kilobytes. Per token. So an 8K-token conversation holds 2.6 GB of KV — about 170 concurrent chats on our 440 GB budget. A 32K chat holds 10 GB → 42 concurrent. A 100K chat holds 32 GB → 13 concurrent, one user's cache now approaching a quarter of the model's weights.
Say that at the whiteboard and watch the interviewer sit up. Context length is a concurrency multiplier in reverse. Ten times the context is roughly a tenth of the users on the same hardware, and no clever code changes that arithmetic. (Grouped-query attention is the trick that keeps that number survivable: instead of every attention head storing its own K and V, groups of query heads share one set. With plain multi-head attention — 64 KV heads, not 8 — KV per token would be 2.5 MB, capping the node near 20 users. GQA is an 8× concurrency win, which is why every serious serving model uses it.)
Committed assumptions: 2M daily users × 6 messages = 12M requests/day ≈ 140 rps average, peak 3× = 420 rps, averaging 1,500 input / 400 output tokens.
Three decisions those numbers just made, which is the only reason to have done them:
POST /v1/chat/completions {model, messages[], max_tokens, stream:true}
headers: Idempotency-Key: <request_id>, Authorization: <tenant key>
→ 200 text/event-stream
data: {"delta":"Quantum"}
data: {"delta":" tunnelling"}
data: {"finish_reason":"stop","usage":{...}}
GET /v1/requests/{request_id} → status + text so far (reconnect / resume)
POST /v1/batch {requests[]} → job_id (the cheap tier)
There's almost no persistent data model, which is itself the senior observation: this system is nearly stateless across requests and enormously stateful within one. We store a request record (id, tenant, token counts, status, partial output — for resume and billing) and per-tenant quota counters; everything else lives in GPU and host memory as KV blocks and prefix cache, and conversation history belongs to the client. That's deliberate: any replica can serve any request, which makes routing a free choice instead of a constraint.
Request in: the gateway authenticates, counts prompt tokens, checks the tenant's tokens-per-minute budget and concurrency cap. The prompt classifier runs (small model, tens of milliseconds, charged to TTFT). The request enters a priority queue by class. The router hashes the prompt prefix, finds the node already holding that conversation's KV blocks, and checks its KV headroom. The node's scheduler admits only if it can reserve blocks for prompt plus expected output. Prefill runs; the first token emerges.
Tokens out: every decode step, the node produces one token for each of ~170 in-flight sequences at once. Tokens land in a small per-request buffer; the streaming layer drains buffers into server-sent events (SSE — one long-lived HTTP response the server keeps writing into); the output classifier inspects text as it flows. Finished sequences release their KV blocks immediately, which is what lets a queued request start on the very next step. Thumb to token to screen, loop closed.
Every extra request you admit buys throughput and costs KV memory — and KV memory is finite, unswappable, and grows every decode step for every user. That makes this a live allocation problem with no good failure mode: too small a batch and you burn money reading weights for nobody; too big and either ITL blows past SLO (throughput up, goodput zero) or you run out of KV mid-generation and the process dies, taking 170 strangers' conversations with it. Worse, TTFT and tokens/sec pull opposite ways — what improves one degrades the other. Everything else in this chapter is a technique for surviving that single trade. Name it out loud and spend your time here.
Static batching is what you write first: collect 32 requests, run them together, return 32 responses, repeat. It has a fatal shape here. Generations differ wildly in length — one user wants yes/no, another wants an essay — and the batch can't return anything until the longest finishes. A 20-token answer sits completed but unreturned while a 2,000-token answer grinds on and its neighbours' slots compute padding. That's the convoy problem, and at realistic length distributions it wastes most of your fleet.
Continuous batching (in-flight batching; introduced as iteration-level scheduling in the Orca paper, OSDI 2022) fixes it with one idea: the scheduling unit is one decode step, not one request. Between every step the scheduler runs — finished sequences leave and free their KV, waiting requests join. The batch becomes a rolling membership instead of a cohort.
Rolling membership needs rolling memory. The old approach reserved, per sequence, one contiguous chunk sized to its maximum possible length. Ask for max_tokens=4096 and generate 200? You reserved 1.3 GB and used 65 MB — times every request — and the gaps left behind are the wrong sizes for new arrivals. Classic external fragmentation.
Operating systems solved this in the 1960s. A program doesn't get one contiguous slab of RAM sized to its worst case; it gets fixed-size pages scattered anywhere in physical memory, tied together by a page table mapping "the program's page 7" to "physical frame 913". Nothing is contiguous, nothing is reserved in advance, and two programs can share a page by pointing at the same frame. PagedAttention is exactly this for KV cache: fixed-size blocks (say 16 tokens), a per-sequence block table, blocks allocated one at a time as tokens appear. Waste is bounded by one partly-filled block per sequence.
In 2023 a Berkeley group (Kwon et al.) measured what was actually happening to GPU memory in production LLM servers and found something embarrassing: existing systems were wasting 60–80% of KV cache memory to fragmentation and over-reservation. Not on the model, not on compute — on empty space held "just in case" a sequence ran to its maximum length. They built vLLM around PagedAttention, and their SOSP 2023 paper reports memory waste falling to under 4% and throughput improving 2–4× at the same latency versus the then state-of-the-art FasterTransformer and Orca. The launch benchmarks were louder still: 14–24× the throughput of stock HuggingFace Transformers, and up to 3.5× HuggingFace's Text Generation Inference. The lesson is a systems lesson, not an ML one — the bottleneck was never the math, it was an allocator making a bad assumption. And once blocks are shared by reference count, sequences with an identical prefix share physical memory for free, which is the feature deep dive 2 is built on.
The scheduler's contract in one sentence: never begin a generation whose memory you cannot finish. Before admitting, reserve KV blocks for the prompt plus a forecast of the output. No budget, no admission — the request waits with an honest queue position, or gets a 429 with Retry-After. Never an accepted request that dies at token 300.
Why so hard-line? GPU out-of-memory is not a per-request failure. The allocation fails inside the model process, and that process is serving the whole batch — one bad admission kills 170 live conversations. Same stance as p3c14: decide the failure direction before 3am. Reject cleanly at the door rather than fail mid-answer.
But pure reservation is too conservative, since max_tokens is an upper bound most requests never approach. So I overcommit deliberately: admit on a forecast (a running percentile of real output lengths, not the declared max), hold a reserve pool, and accept that the bet occasionally loses. When it does, preempt — and how is a real decision:
Commitment: recompute below ~16K tokens (the overwhelming majority), swap above it. Victim: the most recently admitted, lowest-priority request — least sunk cost — with a cap on preemptions per request so nothing starves. This is partly what the batch tier is for: batch-class requests are freely preemptible, which lets me hold headroom for interactive traffic without paying for idle GPUs. The cheap tier is how you afford the fast tier's latency budget.
Here's the most confusing page in this system. Throughput is flat and healthy. ITL p99 has doubled. Nothing deployed. What happened?
Someone sent a 100K-token prompt. Prefilling it takes seconds of solid compute, and during those seconds no decode steps run — so all 170 people mid-sentence stall. One request, a hundred and seventy victims. It's p2c06's hot partition in new clothes: one unit of work costing 100× the median, landing on shared capacity.
Back to the estimate: prefill is half the bill, and multi-turn chat resends the whole conversation every turn. Turn 6 re-prefills turns 1–5 that we computed ten seconds ago. With paged KV those blocks may still be resident, and blocks holding identical tokens are byte-identical — so reuse them. Hash the prompt block by block, keep a prefix cache from block hash to physical block, and on a hit skip prefill for the matched prefix entirely. A sixth-turn message sharing 80% of its prompt costs 20% of the prefill: the difference between a 900 ms TTFT and a 200 ms one, and roughly a third off the fleet.
The catch is routing. That cache lives in one node's memory, and least-loaded routing sends turn 6 to whichever node is quietest, throwing the cache away. So the router must be prefix-aware: hash the conversation prefix, route consistently to the node holding it. This is p1c05's sticky-routing trade-off inverted — there stickiness was a liability that ruined even load distribution; here it's the point, and even distribution is what we give up. Temper it with a load-aware override: if the target node's KV occupancy is above ~85% or its queue is long, route elsewhere and eat the prefill. Affinity is a preference, never an obligation.
One rule that is not optional: key the prefix cache by tenant. Cross-tenant block sharing leaks — a suspiciously fast TTFT tells you someone else recently sent that exact prefix, which is a real side channel for anything sensitive in a system prompt.
Speculative decoding. A small draft model guesses the next 4–5 tokens; the big model verifies all of them in a single forward pass, because verification is parallel like prefill. Every guess it agrees with is free, the first disagreement discards the rest, and done properly the output distribution is identical to the big model's. Pure latency win — at low batch. At high batch the math units are already busy, so speculation spends compute you were using productively and can reduce throughput. Commitment: on for the low-latency tier and off-peak troughs, off above a batch-size threshold. A dial, not a setting.
Quantization. Eight-bit weights halve both the memory they occupy (140 GB → 70 GB, freeing 70 GB for KV) and the bandwidth decode must read, so decode roughly doubles in speed; four-bit doubles that again at a bigger quality cost. The underrated one: quantize the KV cache itself — 8-bit KV halves 320 KB/token and doubles concurrency outright. Stance per tier: 8-bit weights and 8-bit KV by default, full 16-bit for a "high-fidelity" tier customers opt into, 4-bit only for draft models. Every precision change ships behind an eval gate, never on vibes — this is the one knob that degrades the product invisibly.
When the model doesn't fit one GPU. Tensor parallelism splits every layer's matrices across GPUs so they work on each token together, but at every layer all eight GPUs must stop and sum their partial results together (an all-reduce) before anyone can continue — so it only makes sense over a fast intra-node link like NVLink. It keeps single-request latency low: TP=8 within a node. Pipeline parallelism splits by layers across machines (box 1 runs layers 1–40, box 2 runs 41–80). Communication is just boundary activations, so ordinary networking survives it, but it creates bubbles: stage 1 idles while stage 2 works, unless many micro-batches are in flight. Better throughput, worse latency. Commitment: tensor-parallel inside a node, replicate whole nodes, reach for pipeline parallel only if the model stops fitting in one node.
Streaming. Server-sent events over one long-lived HTTP response — simple, one-directional, proxy-friendly, reconnects natively; WebSockets buy bidirectionality we don't need. The interesting failure is a slow client: someone on a train whose socket drains at a trickle. If the GPU slot waits on that socket, one bad connection holds 2.6 GB of KV hostage. So decouple — the model writes into a small bounded per-request buffer and the HTTP layer drains it independently, so generation runs to completion at full speed regardless. If the buffer fills we stop sending, not generating: finish, persist, let the client resume by request id. That's p2c07's backpressure at the right layer — never let the expensive stage block on the cheap one.
Retries, honestly. If a node dies at token 200, that KV is gone and there is no resume — the state was ephemeral by construction. Idempotency here means re-running produces one answer and one bill, not that we continue mid-sentence. Every request carries a request id, the gateway dedups on it, and a retry either attaches to the in-flight generation, replays a persisted completed stream, or restarts a dead one — usage written once, keyed by that id (p2c02, same double-charge discipline as p3c14). Restarting is non-deterministic at nonzero temperature, so surface it as a regeneration rather than pretending nothing happened.
Safety in the path. Two checkpoints with latency budgets. The prompt classifier runs before prefill — small model, tens of milliseconds, charged straight to TTFT. Output moderation is harder because output arrives progressively: checking only at the end defeats streaming, checking every token is too expensive. So check in windows, every ~50 tokens or at sentence boundaries, over the accumulated text.
That leaves a genuine product choice, and I'd commit rather than hedge. Either hold back (emit only text that already passed a check, adding delay to the visible stream) or emit then revoke (send an explicit event telling the client to remove what it displayed). Holding back costs perceived speed; revoking means a user occasionally watches text vanish, which feels worse — worse still if they screenshotted it. Commitment: hold back a small window for the highest-risk categories, revoke-and-explain for the rest, with a first-class revoked event designed in from day one, because retrofitting that into a client is miserable. Two things people forget: the classifiers are models too, with their own GPUs and capacity plan, and the fail-open/fail-closed rule must be pre-decided — closed for the highest-risk categories, open with loud logging for the rest, because globally fail-closed turns a classifier hiccup into a total outage.
Fairness and rate limits. Requests-per-minute is the wrong unit when one request can be 50 tokens or 200,000. Limit on tokens per minute, counting output several times heavier than input because decode is the scarce resource; charge the estimate at admission, reconcile on completion. Add a per-tenant concurrency cap, because concurrency is what consumes KV — a tenant with a modest token budget opening 500 simultaneous streams still eats the node. Enforcement is p3c10's machinery: local token buckets per gateway with async reconciliation, accepting small overshoot rather than a Redis round-trip in front of every request.
Semantic caching, and why I mostly don't. An exact-match cache keyed on a hash of (model, prompt, sampling params) is trivial and safe, and in chat it almost never hits — real conversations are unique. It earns its keep for programmatic traffic: classification calls, fixed product prompts, temperature-zero requests. Fuzzy semantic caching — embed the prompt, serve a near neighbour's answer — I decline for the general product, with the reason on record: two prompts can be 0.98 cosine-similar and need opposite answers ("is X safe for a child" versus "is X safe for a child with a nut allergy"). A wrong fuzzy hit is a silent correctness bug with no error to alert on, and the prefix cache is where the real reuse lives anyway.
What goes on the wall, unprompted:
Node death mid-stream is frequent at 500 GPUs, and one dead GPU takes its whole tensor-parallel group with it. The answer isn't prevention, it's a clean client contract: request ids, persisted partial output, reconnect to resume or regenerate, billed once. Blast radius is the batch on that node — another argument for batching at the SLO-optimal point rather than the maximum.
Autoscaling, honestly. A new node must pull a container, fetch 140 GB of weights, load them into GPU memory and warm up kernels — minutes. Reactive autoscaling on latency is useless: the node arrives after the spike ends. So: predictive scaling on the daily curve (chat traffic rises and falls with waking hours in each region, which makes it unusually easy to forecast), a warm pool of pre-loaded idle nodes as insurance (5% of a $9M fleet is cheap next to a capacity incident), node-local NVMe weight caching so restarts don't re-download, and for unforecast spikes degrade rather than scale: pause the batch tier, tighten default max_tokens, lengthen the queue, shed with honest 429s.
At 10x? Prefill compute breaks first, because long-context usage grows faster than user count — prompts lengthen as people trust the product. KV memory second, as those longer contexts collapse per-node concurrency. Decode bandwidth last. So: prefix caching harder, 8-bit KV, split prefill and decode into separate pools, and the biggest lever of all — route by difficulty, sending the easy 60–70% to a much smaller model and reserving the 70B for what needs it. That last one is a product decision as much as an infrastructure one, and saying so is part of the answer.
Anthropic published an unusually detailed postmortem on three bugs that degraded Claude's output quality in late 2025 — and every one was a serving-infrastructure bug, not a model bug. The first matters most here: a context-window routing error sent short-context requests to servers configured for the 1M-token context window. Because routing was sticky, a request that landed on the wrong server tended to keep landing there — follow-up turns in the same conversation were likely served by the same incorrect server. It began at 0.8% of Sonnet 4 requests and grew: at the worst hour on August 31, 16% of Sonnet 4 requests were affected, and roughly 30% of Claude Code users who made a request in that window had at least one message land on the wrong server type. The second was a TPU misconfiguration where a runtime performance optimization assigned high probability to tokens that should never have been chosen — English prompts answered with Thai or Chinese characters, or code with obvious syntax errors. The third was a latent XLA:TPU compiler miscompilation of the approximate top-k operation used to pick high-probability tokens, triggered by an unrelated deployment. Three lessons for your design. A heterogeneous fleet — nodes configured for different context sizes — is a routing correctness problem, not just a capacity problem, and the sticky routing you added for prefix-cache hits will faithfully pin users to a broken node. Optimizations in the sampling path can change outputs, so they need output-quality monitoring, not just latency monitoring. And "the answers feel worse" is invisible to every dashboard in the list above unless you deliberately build quality telemetry beside the latency telemetry.
They'll let you draw boxes for five minutes, then aim everything at memory and scheduling. Expect close to these words:
| Decision | Chosen over | What it costs me |
|---|---|---|
| Continuous batching with paged KV blocks | Static batching, contiguous reserved KV | A far more complex scheduler and allocator, per-step scheduling overhead, block tables to keep correct under preemption |
| Admission control on a KV budget, with deliberate overcommit | Accept everything and hope | Visible queueing and 429s at peak instead of silent degradation; forecast errors force preemptions |
| Recompute on preempt below 16K, swap above | Never preempt; or always swap | Preempted users feel a hiccup, and recompute burns that request's prefill twice |
| Chunked prefill | Prefill-to-completion, first come first served | Long prompts get worse TTFT and prefill runs slightly less efficiently — paid so 170 streams don't stall |
| Prefix-aware sticky routing, tenant-scoped | Pure least-loaded routing | Uneven node load, hot replicas, and a correctness hazard: a bad node stays bad for the users pinned to it |
| 8-bit weights and 8-bit KV as the default tier | 16-bit everywhere | A small, real quality cost only an eval suite can see — the one knob that degrades the product invisibly |
| Tensor parallel in-node, replicate nodes | Pipeline parallel across nodes | The model must fit one node, and capacity grows in 8-GPU units |
| Predictive scaling plus a warm pool | Reactive autoscaling on latency | Paying for idle GPUs, and unforecast spikes handled by shedding rather than growing |
A junior asks: "It's just a model behind an API — why can't we scale it like our normal web service?" Explain in five or six sentences.
A normal web request is stateless and takes milliseconds; an LLM request holds a growing chunk of GPU memory for twenty seconds while it writes one token at a time. That memory is the KV cache — the model's saved attention state — and it grows with every token, so a long conversation holds gigabytes and one node only fits a couple of hundred conversations. Generating a single token means reading the entire model's weights out of GPU memory, which is so wasteful for one user that the only economical approach is serving many users in the same batch, sharing that one read. So "load balancer plus autoscaler" doesn't fit: between every token step, a scheduler has to decide who is in the batch and whether there is enough memory to let anyone new in. And you can't scale out of a spike, because starting a node means loading 140 GB of weights, which takes minutes — so you forecast, keep warm spares, and degrade gracefully instead.
Now add a 1M-token context tier — customers want to paste whole codebases. Redo the arithmetic and say what actually has to change.
KV for one 1M-token request at 320 KB/token is 320 GB — more than double the model's weights, half a node, for one user; shared-node concurrency goes to roughly one. Prefill is worse: 2 × 70e9 × 1e6 ≈ 1.4 × 10^17 FLOPs, tens of seconds of node compute, so TTFT is measured in tens of seconds and the product must show progress or go async. Changes: a dedicated node pool with its own SLOs and pricing (never mix these with interactive chat), 8-bit or 4-bit KV to cut that 320 GB, offloading cold blocks to host memory or NVMe with prefetch, and aggressive prefix caching because these users resend the same codebase every turn. And note the operational hazard from the Anthropic postmortem: once the fleet is heterogeneous by context size, misrouting between pools becomes a correctness bug — route on declared context class and alarm on mismatches.
10x the traffic: 4,200 requests/second at peak. Name what breaks first, in order, and what you do about each.
(1) Prefill compute — 6.3M prompt tokens/second needs ~250 nodes for prefill alone, and it's the half of the bill you can attack without buying hardware: raise prefix-cache hit rate, add cross-node prefix lookup, and split prefill and decode into separately-sized pools so you stop over-provisioning decode memory to buy prefill FLOPs. (2) KV memory — quantize KV to 8-bit for an instant 2× concurrency, tighten default max_tokens so forecasts reserve less, grow the preemptible batch tier. (3) The router becomes its own scale problem: prefix-hash tables plus per-node occupancy at 4,200 rps want a sharded, gossip-updated view rather than one central scheduler. (4) The honest answer underneath: at 10× the real move is model tiering — route the easy majority to a small model, keep the 70B for what needs it. Bigger than any scheduling change, and I'd say so before being asked.
Requirement flips: a regulated enterprise customer demands their requests never share a batch with anyone else's. Quantify the cost and propose something you'd actually ship.
Quantify first: batching is where the economics live, so a batch of one runs at roughly 1/170th the token throughput of a shared node — two orders of magnitude more per token. Then get precise about what they actually fear. Sequences in a shared batch don't leak into each other; attention is computed per sequence, and batching is a matrix-shape optimisation, not shared state. The genuine cross-tenant surface is the prefix cache, where shared blocks and TTFT timing can reveal that someone else sent the same prefix — which is why ours is keyed by tenant. So the shippable middle: dedicated nodes (not dedicated batches) on a minimum-commitment contract, so their own traffic batches together and they keep ~90% of the efficiency; tenant-scoped prefix cache and per-tenant KV accounting; a documented isolation model. Committing to the mechanism beats agreeing to the demand.
"Cut serving cost 40%. You may not change the model or degrade the interactive SLO." Where do you go?
Ranked by win per unit of risk. (1) Prefix caching plus prefix-aware routing — prefill is half the compute and chat resends most of its prompt, so a good hit rate takes a large bite with zero quality change. (2) 8-bit KV cache — doubles concurrency per node, attacking the binding constraint, with far less quality impact than weight quantization. (3) Push async traffic into the batch tier, on preemptible capacity, since it is preemptible by design already. (4) Right-size defaults: max_tokens of 4,096 when the p95 answer is 400 tokens makes every admission reserve ten times what it needs. (5) Speculative decoding off-peak, when batches are small and the FLOPs really are free. What I would not do is raise batch size past the ITL SLO — that trades goodput for a throughput number while looking good on a dashboard.
Final round. The interviewer smiles and delivers the 2026 closer: "Our company has a hundred million internal documents. Design the system that lets an AI assistant answer questions about them — accurately, with citations, in a couple of seconds."
Watch the trap spring. Mid-level candidates hear "AI assistant" and start comparing LLMs. That's the one part of this system you rent. The part you have to design is everything in front of the model: turning a billion pieces of text into vectors, searching them in under a hundred milliseconds, keeping them fresh as documents change, and making sure a sales rep never retrieves the CEO's M&A folder. This is a full-blooded distributed systems problem wearing a GenAI costume — sharding, caching, migrations, multi-tenancy, the whole book.
And the stakes are nastier than ordinary search. A search engine with bad ranking shows you bad links; you notice and rephrase. A RAG system — retrieval-augmented generation, where the model answers from documents you hand it — doesn't fail loudly. Feed the LLM the wrong paragraphs and it will compose a fluent, confident, cited, wrong answer. Retrieval quality gates the model's honesty. That single sentence is the frame for the next forty-five minutes, and I'll say it to the interviewer in the first two.
Functional: semantic search over all tenant documents — meaning search by meaning, so "how do I expense a flight?" finds the travel policy even though it never contains the word "expense"; a Q&A endpoint that retrieves relevant passages and feeds an LLM to compose an answer with citations back to source documents; multi-tenant — we're a B2B product, thousands of companies, and access control lists (ACLs) inside each tenant, because HR docs aren't for everyone; continuous ingestion — new and edited documents searchable within minutes, not after a nightly build. Out of scope: the LLM serving stack itself (that was last chapter, p3c23 — we call it over an API) and training our own embedding model.
Non-functional, with numbers I'll defend:
An embedding model is a neural network that reads a piece of text and outputs a vector — a list of, say, 768 floating-point numbers. The magic property: texts with similar meaning land near each other in this 768-dimensional space. "How do I get reimbursed for a flight?" and "Travel expense policy" end up close; "flight" and "Apache Flink" end up far apart. Search becomes geometry: embed the query, find the nearest document vectors, return those documents. Distance is usually measured by cosine similarity or dot product — for this chapter, just read "nearest = most similar in meaning."
Now the trap, stated up front because it's the most common silent failure in real RAG systems: the corpus and the queries must be embedded by the exact same model version. Each model defines its own private geometry. Vectors from model v7 and model v8 are maps of two different countries — comparing a v8 query against v7 documents doesn't throw an error, doesn't crash, doesn't log a warning. Every distance is simply meaningless, and search quietly returns garbage. So: every vector in my system is stamped with model_version, the query path asserts its version matches the index it's searching, and a mismatch is a hard, loud error. And upgrading the embedding model means re-embedding the entire corpus — a full migration we'll walk through later, because it's this problem's version of the p2c11 exam.
Do the one multiplication that matters, out loud:
1 billion vectors × 768 dimensions × 4 bytes (fp32) = 3,072 bytes each ≈ 3 TB of raw vectors.
That number makes three decisions for me in ten seconds. First, no single machine — 3 TB of vectors (plus index structure on top) must be sharded. Second, compression is mandatory — RAM at 3 TB+ across a fleet, replicated 3×, is real money; we'll squeeze vectors ~30× with quantization. Third, exact search is off the table at this size — comparing a query against a billion vectors is ~1.5 trillion floating-point operations per query; at 5,000 queries/sec that's a physics problem, not an engineering problem. We need approximate nearest neighbor search — ANN — which finds almost certainly the nearest vectors while touching only a tiny fraction of them. ANN at scale is this problem's crux, and I'll spend my longest deep dive there.
But first, the honesty check that echoes p0c01's over-engineering warning. Flip the numbers small: at 1–10 million vectors, brute force — literally computing the exact distance to every vector — is the right answer. 1M vectors × 3 KB = 3 GB: fits in one machine's RAM with room to spare. With SIMD instructions — the CPU feature that multiplies eight or sixteen numbers in one instruction instead of one at a time — or a single GPU, an exact scan over a million vectors takes single-digit milliseconds, batched. You get 100% recall, zero index build time, zero tuning knobs, and inserts are just appends. A startup with 2M vectors that deploys a sharded ANN cluster has bought itself an on-call rotation for nothing — pgvector inside the Postgres you already run, or a flat FAISS index, ends the conversation. Say this in the interview before building the big machine: "below ~10M vectors I wouldn't build any of what follows." That sentence is a senior signal all by itself.
Two more numbers. Ingest: 1B chunks at ~250 tokens each is ~250 billion tokens through the embedding model — at realistic batched-GPU throughput, that's a day-or-more job on a serious GPU pool. Conclusion: full re-embeds are projects, so the pipeline must support incremental updates as the normal case. Query side: 5,000/sec × ~10 candidate lookups each is trivial for the document store, but the reranking stage (coming later) runs a neural model per candidate — that's the compute hotspot to budget, not the key-value reads.
POST /v1/query — {tenant_id, user_id, query_text, top_k, filters} → ranked passages with scores and doc references. This is the retrieval endpoint; the answer service calls it.POST /v1/answer — same inputs → streamed LLM answer plus a citations array mapping each claim to chunk IDs.PUT /v1/documents/{doc_id} / DELETE ... — upsert or remove a document; kicks the ingest pipeline.The unit of storage and retrieval is the chunk, not the document: {chunk_id, doc_id, tenant_id, acl_refs, text, title, section_path, model_version, updated_at, vector}. Vectors live in the ANN index (sharded); full chunk text and metadata live in a boring replicated document store keyed by chunk_id; the ANN index stores only vector + chunk_id + the filterable attributes (tenant_id, acl_refs, updated_at). Keep the fat text out of the RAM-expensive index.
Write path (ingest): a document arrives or changes → the chunker splits it → the embedder batches chunks through the model, stamping model_version → vectors upsert into the ANN shards (routed by hash of chunk_id within the tenant's shard set), text into the doc store, tombstones for deleted chunks. It's a queue-driven pipeline (p1c11): bursty ingest never touches query latency, and a poison document can't wedge anything but its own retry.
Read path (query): embed the query with the same model → fire two searches in parallel — keyword search (BM25) and vector search (ANN) — merge the candidate lists, rerank the top 100 with a heavier model, assemble the top passages into a prompt with citation anchors, and hand it to the LLM. Every stage after "embed" is fanned out across shards and merged, classic scatter-gather.
You can't embed a 40-page document as one vector — one vector can't represent forty pages of distinct meanings, and you can't stuff 40 pages into a prompt anyway. So we split. How you split matters more than most teams expect, because the chunk is what gets retrieved and what the LLM reads. My commitments: chunks of roughly 200–500 tokens — big enough to carry a complete thought, small enough that one vector represents one topic. 10–20% overlap between neighboring chunks, so a sentence living at a boundary exists whole in at least one chunk. Structure-aware splitting: break at headings, paragraphs, and list boundaries — never mid-sentence; a chunk that starts with "…and therefore the limit is $50" is a riddle, not a passage. And every chunk carries its context as metadata: document title and section path get prepended to the text before embedding ("Travel Policy > International > Reimbursement: …"), because the sentence "the limit is $50" embeds uselessly without knowing what it's a limit on. Cheap trick, large recall gains, ask me why in the eval section.
This is the hard part the interview question is invented to probe. Exact nearest-neighbor search at 1B vectors is computationally impossible within our latency budget, so every real system approximates — and approximation drags you into a three-way tug-of-war: recall (what fraction of the true nearest neighbors you actually find) vs latency (how much of the index you can afford to touch per query) vs memory (what the index costs to hold). Pull any corner and the other two move. Every index family — graphs, clusters, disk-based — is just a different stance in this triangle, and the senior move is to name the triangle, pick a stance for these numbers, and say what it costs.
Option one: HNSW — Hierarchical Navigable Small World graphs, the index inside most modern vector databases. Every vector is a node connected to a few dozen near neighbors. Search greedily walks the graph: from an entry point, repeatedly hop to whichever neighbor is closest to the query, until no neighbor improves. The "hierarchical" part fixes greedy search's slowness: stack multiple layers, where the top layer holds a tiny sample of nodes with long-range links and each lower layer gets denser, down to layer 0 which holds everything.
Driving across a country to a specific house. You don't take neighborhood streets the whole way — you get on the highway (top layer: few exits, huge hops), exit near the right city (drop a layer: avenues, medium hops), then wind through local streets (layer 0: every house exists here) to the door. HNSW searches exactly like this: coarse fast hops first, precise short hops last — a skip list's trick, applied to a graph. Where the analogy breaks: roads are 2D and reliable; in 768 dimensions the "streets" are statistical, so occasionally greedy driving parks you at the wrong house — that's the "approximate" in ANN.
HNSW's report card. Recall/latency: excellent — 95–99% recall at sub-10ms per shard, tunable at query time with one knob (ef_search: how wide a frontier the search keeps; higher = better recall, slower). Inserts: incremental and cheap — a new vector links itself in immediately, which is exactly what our 5-minute freshness demands. The bills: RAM-hungry — the graph needs the vectors plus the links, roughly 100–250 bytes per node at typical settings (32 neighbors at layer 0, four bytes an ID), all hot in memory; builds are slow (inserting a billion nodes one by one takes serious time — another reason to shard); deletes are awkward — you tombstone the node (mark it dead so queries skip it, leaving the bytes in place) and physically remove it in a later compaction.
HNSW comes from Yury Malkov and Dmitry Yashunin, who published it in 2016 building on their earlier "navigable small world" graph work — research done at institutes in Nizhny Novgorod, Russia, far from any hyperscaler. The paper's core trick is explicitly borrowed from the skip list: keep a hierarchy of ever-sparser levels so you can move coarsely first and precisely last. It spent a couple of years as an academic curiosity, then the embedding era arrived, and its recall-latency numbers on standard benchmarks beat everything practical. Within a few years it became the default index in nearly every vector database — Qdrant, Weaviate, Milvus — was added to Apache Lucene (and therefore Elasticsearch and OpenSearch), and landed in pgvector so a plain Postgres table could do it. Two researchers, one data structure, the retrieval backbone of the GenAI boom. When your interviewer asks "why HNSW?", "it won the benchmarks so thoroughly the whole industry converged on it" is a true and complete answer.
Option two: IVF — inverted file index. Run k-means over the corpus — the clustering algorithm that groups points around a fixed number of centers and moves each center to the middle of its group until it stops moving — to produce, say, 64K cluster centroids (the "coarse quantizer"). Each vector is filed into its nearest centroid's list. At query time: compare the query to the 64K centroids (cheap), pick the nearest nprobe clusters — say 16 — and exhaustively scan only those lists. You've searched 16/64,000ths of the corpus. IVF is cheaper in memory than HNSW (lists, no graph), builds fast, and nprobe is a clean recall-vs-latency dial. Its weakness is the recall cliff at cluster boundaries: if the true nearest neighbor sits just across the border in a cluster you didn't probe, it is invisible — no amount of scanning the wrong lists finds it. And because centroids are computed once, a drifting corpus (new product names, new topics) slowly degrades the clustering until you re-train and rebuild — a scheduled operational burden HNSW doesn't have.
Option three: disk-based graphs (the DiskANN family, from Microsoft Research, 2019). Keep compressed vectors in RAM for navigation but the full graph and exact vectors on NVMe SSD — the paper reported a billion points indexed and served from a single 64 GB workstation with a cheap SSD, over 5,000 queries/sec at under 3 ms mean latency and 95%+ recall@1. It trades tail latency and much slower index builds for fitting 5–10× more vectors per machine than an in-RAM graph. At our scale it's the credible cost-saving alternative, and I name it so the interviewer knows I know — then I commit elsewhere.
The commit. My sizing rule, stated as policy: under ~10M vectors — brute force, no index. 10M–100M — HNSW in RAM on one box plus replicas. Around 1B — shard, compress, two-stage. For this system: 32 shards, ~30M vectors each, HNSW per shard, built over compressed vectors, with exact re-ranking on top (compression next paragraph). Why HNSW over IVF here: our 5-minute freshness requirement makes HNSW's incremental inserts decisive — IVF's centroid drift and periodic retrain-rebuild cycle is an operational tax that fights our freshness story. Why not DiskANN day one: RAM-resident HNSW with quantization already hits budget, and p99 predictability is easier to defend in RAM; if the fleet bill or a 10× corpus bites, DiskANN is my named escape hatch. Each shard gets 2 replicas — that's availability and read throughput, and the arithmetic is worth saying out loud: scatter-gather means every query touches every shard in the tenant's set, so each shard sees the full 5,000 QPS, not 5,000 ÷ 32. Three copies per shard is what turns that into ~1,700 QPS a box.
Product quantization (PQ) compresses each vector ~30× by describing it with a small vocabulary. Chop the 768-dim vector into 96 sub-vectors of 8 dimensions. For each of the 96 positions, learn a "codebook" of 256 representative patterns (via k-means). Now store each sub-vector as just its closest pattern's ID — one byte. The whole vector becomes 96 bytes instead of 3,072: 32× smaller, and 1B vectors drop from ~3 TB to ~96 GB — suddenly 32 shards of ~3 GB each plus graph links, comfortable RAM territory. Distances are computed against the codes via lookup tables, fast and SIMD-friendly.
A paint store doesn't describe a wall color by its full light spectrum; it says "Sherwin-Williams 6244." A small swatch book of named colors stands in for infinite possible colors. PQ gives every 8-dimensional slice of your vector a swatch book of 256 entries and records only the swatch numbers. You lose nuance — two slightly different walls may map to the same swatch — which is exactly why PQ costs recall: nearby-but-distinct vectors can become indistinguishable after compression.
That recall loss is real — and here's the standard two-stage trick that claws it back. Stage one: search the compressed index for the top ~100 candidates. Compressed distances are slightly wrong, so the true best answer is probably in that 100 but maybe not ranked first. Stage two: fetch the exact full-precision vectors for just those 100 candidates (from SSD — 100 random reads is nothing), recompute exact distances, re-sort, return the top 10. The cheap fuzzy index finds the right neighborhood; the exact re-rank picks the right house. You pay ~100 exact distance computations instead of a billion, and recover most of the recall PQ cost you. This candidates-then-refine shape is the same pattern the whole read path uses — you're about to see it again, one level up.
Product quantization isn't GenAI-era tech — it comes from a 2011 paper by Hervé Jégou, Matthijs Douze, and Cordelia Schmid at INRIA, solving image search when "web-scale" meant matching billions of image descriptors on 2011 hardware. Jégou and Douze later landed at Meta AI, and in 2017 the same lineage shipped as FAISS — Facebook AI Similarity Search — the open-source library that first made billion-scale similarity search on GPUs a download instead of a research project; the companion paper was literally titled "Billion-scale similarity search with GPUs." FAISS's workhorse configuration is exactly the stack this chapter just built: IVF or graph navigation over PQ-compressed vectors, then exact re-ranking of a short candidate list. When the LLM wave hit in 2022–2023 and every company suddenly needed vector search, a decade of this quiet infrastructure was sitting there ready — FAISS became the embedding index inside countless RAG stacks, and its ideas (PQ codes, coarse quantizers, two-stage refinement) are load-bearing in nearly every vector database you can name today.
| Index | Recall / latency | Memory | Freshness | Pick it when |
|---|---|---|---|---|
| Brute force (flat) | 100% recall; fine to ~1–10M vectors | Raw vectors only | Perfect — append and go | Small corpus. The honest default; anything more is over-engineering. |
| HNSW | 95–99% recall, sub-10ms/shard; ef_search dial | Highest: vectors + graph links, all in RAM | Incremental inserts — best of the ANN family | Latency-critical, update-heavy, RAM affordable. Our commit (over PQ codes + exact re-rank). |
| IVF (+PQ) | Good, but recall cliff at unprobed clusters; nprobe dial | Low — lists, no graph | Inserts OK, but centroid drift forces periodic retrain/rebuild | Batch-built corpora, tight memory, GPU scan (FAISS-style). |
| DiskANN-class | ~95% recall, a few ms — from SSD | Lowest RAM: codes in memory, graph + exact vectors on NVMe | Builds slow; updates batched | Billions of vectors, hardware budget matters more than p99 tail. Our named 10× escape hatch. |
Pure vector search has a blind spot that bites in production: exact tokens. A support engineer searches ERR_CONN_RESET_4012. That string is an arbitrary identifier — the embedding model has no semantic hook for it, so dense retrieval happily returns passages about "connection problems" in general while the one runbook containing that exact error code sits unretrieved. Meanwhile classic keyword search — BM25 over an inverted index (p3c19) — nails exact tokens, IDs, SKUs, function names, people's names… and is hopeless at paraphrase: "why can't my server reach the database at night" shares almost no words with "scheduled firewall rule blocks egress after 22:00," which dense retrieval catches easily. Each engine covers the other's blind spot. So run both, in parallel, and merge.
Merging has a snag: BM25 scores and cosine similarities live on incomparable scales — you can't just sort the union by score. Reciprocal rank fusion (RRF) sidesteps scores entirely and uses only ranks: each document's fused score is the sum over lists of 1 / (60 + rank). Appear near the top of either list and you score well; appear in both and you score better; the constant 60 damps the difference between rank 1 and rank 3 so neither engine dominates. It's embarrassingly simple, needs no tuning or calibration, and is the industry-standard default — Elasticsearch, OpenSearch, and most vector databases ship exactly this.
Then the final quality stage: the fused top ~100 goes through a cross-encoder reranker. Our retrieval embeddings come from a bi-encoder — query and document embedded separately, meeting only at a dot product; that separation is what makes indexing possible, and it's also a lossy summary. A cross-encoder reads the query and the candidate chunk together through one transformer, attending token-by-token across both — dramatically more accurate, and far too slow to run against a billion documents. Against 100? A batched GPU call, ~30–40 ms. Candidates-then-refine again: the index finds 100 plausible passages fast, the expensive model picks the best 10 carefully. In quality evaluations this stage is routinely the single biggest win in the whole pipeline — it is worth its latency, and the eval section is where that stops being a claim and becomes a number.
The latency budget, out loud, because the interviewer is checking it fits inside 100 ms: embed query ~8 ms → BM25 and ANN in parallel, ~20 ms for the slower → RRF ~0 → fetch 100 chunk texts ~8 ms → cross-encoder ~35 ms → total ≈ 70–80 ms p50, with p95 inside budget. The reranker is the biggest line item and the first thing we shed under overload (ops section).
Normal freshness is already solved by the index commit: HNSW takes incremental upserts, so edit-to-retrievable is just pipeline lag — seconds of queueing plus an embed call, comfortably inside 5 minutes. Deletes tombstone immediately (filtered at query time) and physically compact in background segment rebuilds, during which the old segment serves reads and swaps atomically — the reader never sees a half-built index, only a bounded-stale one.
Embedding model migration is the event that breaks naive designs, because of the skew trap from minute six: v8 queries cannot search v7 vectors, so there is no such thing as upgrading in place, chunk by chunk. It's a dual-index shadow migration — p2c11's playbook, almost verbatim: (1) stand up an empty v8 index alongside serving v7; (2) dual-write: from now on every upsert embeds twice, into both; (3) backfill: stream the whole corpus through the v8 embedder into the shadow — the day-scale GPU job from our estimate, rate-limited so it doesn't starve live ingest; (4) shadow-read: mirror a sample of production queries to v8 and compare against the eval suite — this gate is why the eval section exists; (5) cut over tenant by tenant behind a flag, queries now embedding with v8; (6) keep v7 warm for a rollback window, then delete. Cost: double embedding compute and double index RAM for the migration's duration. That's not waste — that's what "upgrade the embedding model" actually costs, and saying the price unprompted is the senior signal.
Multi-tenancy and ACLs. Here's the trap to teach concretely. Naive design: search first, filter after. An engineer at tenant Acme asks about "Project Falcon." Falcon is mostly discussed in an executives-only strategy folder — so the 100 nearest chunks are all documents this user can't read. Post-filter deletes all 100, and the user gets zero results — while a perfectly readable HR announcement mentioning Falcon sits at rank 140, never fetched. Post-filtering starves top-K whenever permissions correlate with topic, which in a company corpus is always. The fix: filter during search — push the predicate into the index, so HNSW's graph walk skips non-matching nodes but keeps walking until it accumulates K visible results. Every serious vector database supports filtered search natively for exactly this reason.
Tenant placement, committed: big tenants (millions of chunks) get their own namespace — dedicated shard set, physical isolation, per-tenant index tuning, trivially satisfies "never." The long tail of small tenants shares sharded indexes with a mandatory, non-optional tenant_id filter compiled into every query at the gateway (defense in depth: the doc store checks it again at fetch). Pure per-tenant indexes for 10,000 tenants would mean 10,000 tiny fragmented indexes — operationally silly; pure sharing puts noisy neighbors in one blast radius. Split the difference where the numbers say to, and add per-tenant rate limits (p1c12) so one tenant's crawler can't eat the fleet's rerank GPUs.
Here's a sentence most candidates never say, and interviewers remember when someone does: "none of my retrieval choices are defensible without an evaluation set." Every knob in this system — ef_search, chunk size, RRF vs learned fusion, rerank depth, PQ bits — trades quality against cost, and without measurement you're tuning blind. So, minimum viable eval: a few hundred real user queries, each labeled by humans with the chunks that actually answer it — the golden set. Metric: recall@10 — for what fraction of queries does the labeled chunk appear in our final top ten? Run it like a CI suite on every index config change, model upgrade, and chunker tweak. This is the gate in the migration's shadow-read step, the proof behind "the reranker is worth 35 ms," and the reason the metadata-prepending chunking trick earned its place — each of those claims becomes a number on this suite instead of a vibe.
Answer-level quality needs one more layer: faithfulness — is every claim in the generated answer actually supported by a cited chunk? Check it by sampling production answers and asking a cheap LLM judge "does this cited passage support this sentence?" — imperfect, but it trends, and trends page people. And when a bad answer is reported, the debugging split that keeps teams sane: look at the retrieved context first. Was the right chunk in the prompt? No → retrieval bug: embedding, chunking, filters, index recall — this pipeline's problem. Yes, and the model ignored or contradicted it → generation bug: prompt assembly or the LLM — last chapter's problem. One question splits the debugging space in half, and it's the first question a senior asks.
LinkedIn's customer-service teams used a RAG assistant to answer support questions from their internal knowledge base, and in a 2024 paper they documented what moved the needle — and it wasn't a bigger LLM. Their diagnosis was retrieval-side: naive chunk retrieval kept losing the structure connecting related issues, so the model answered from fragments. They rebuilt retrieval around a knowledge-graph representation of historical support tickets — preserving the relationships between an issue, its causes, and its resolutions instead of shredding everything into isolated text chunks — and evaluated end to end against their existing baseline. The published result: median per-issue resolution time dropped by 28.6%. The lesson this chapter cares about: they found the win by treating retrieval quality as the measurable bottleneck of answer quality, running the comparison, and shipping what the numbers picked — the eval-driven loop, working exactly as designed.
The cost triangle closes the loop: this system spends money in three places — embedding compute (ingest and re-embeds), index memory (RAM fleets), and LLM tokens (every query, forever, scaling with how much context you stuff). The tempting hack is to skip retrieval quality and shovel 50 chunks into a giant context window "to be safe." That's the worst corner of the triangle: maximum token cost, and measurably worse answers, since irrelevant context actively distracts models. Better retrieval means sending 5 right chunks instead of 50 hopeful ones — cheaper per query and more faithful. Precision is the one investment that pays on both axes; the eval suite is what lets you buy it deliberately. Also assemble the prompt with intent: dedup near-identical chunks (don't spend budget on five copies of the same boilerplate), prefer diversity across documents, and attach citation anchors ([1], [2] → chunk IDs) so the LLM can ground each claim.
How interviewers probe this design once it's on the board — each question is checking whether the machine in your head actually runs:
ef_search / rerank depth / chunking, priced in latency and GPU, chosen via the eval suite — not "use a better model."| Decision | Chose | Over | The cost I accept |
|---|---|---|---|
| Index | 32-shard HNSW over PQ codes + exact re-rank | IVF-PQ; DiskANN | Highest RAM of the family; slow builds; tombstone-and-compact deletes |
| Compression | PQ ~32× + two-stage exact re-rank | Raw fp32 in RAM | Recall loss partially recovered; extra SSD fetch + recompute per query |
| Retrieval | Hybrid BM25 + dense, RRF fusion | Dense-only | Two indexes to keep consistent from one ingest pipeline |
| Quality stage | Cross-encoder rerank 100 → 10 | Trust index order | ~35 ms and a GPU pool — the latency budget's biggest line |
| Tenancy | Namespaces for big tenants, filtered shared shards + in-search ACL for the tail | All-shared or all-dedicated | Two placement code paths; filter pushdown must be bulletproof |
| Model upgrades | Dual-index shadow + eval-gated cutover | In-place re-embed | ~2× embed and index cost for the migration window |
| Quality assurance | Golden-set recall@10 as CI + prod canaries + sampled faithfulness judge | Ship and hope | Ongoing labeling effort — the cheapest component in the system |
A junior asks: "Why can't we just compare the query against all billion vectors and get the exact right answer? And what does HNSW actually do instead?" Explain in five or six sentences, including what we gave up.
Comparing against all billion vectors means about 1.5 trillion arithmetic operations per query — at thousands of queries per second, no fleet we can afford keeps up, so exact search dies at this scale. HNSW instead organizes vectors into a graph where each one links to its near neighbors, stacked in layers like highways over streets: the sparse top layer makes huge hops to get near the target fast, and each denser layer below refines the position until the bottom layer walks the last few steps. A search touches a few hundred vectors instead of a billion, which is why it returns in milliseconds. What we gave up is certainty: it's approximate, so maybe 1–5% of the time the true nearest neighbor isn't in our results — we measure that as recall and tune a knob to trade speed for accuracy. We also pay in memory, because the graph and vectors must sit in RAM, which is why we compress vectors and shard across machines. Below ten million vectors or so, none of this is worth it — brute force is exact, simple, and fast enough.
Estimation drill: a corpus of 200M chunks with 1024-dim fp32 embeddings. Raw vector size? After PQ at 128 bytes per vector? What deployment does each number imply?
Raw: 200M × 1024 × 4 B = 4,096 B each ≈ 800 GB — too big for one box in RAM, so raw fp32 means sharding immediately. PQ at 128 B: 200M × 128 B ≈ 25.6 GB — fits in one machine's RAM with the graph on top, so the honest design is one HNSW-over-PQ node plus two replicas and exact re-rank from SSD, no sharding at all. Same corpus, opposite architectures — the compression decision is the topology decision, which is why you do this math before drawing boxes.
A legal-tech startup: 2M documents → ~20M chunks, 50 queries/sec, corpus grows 2% a month. The founder wants "the Pinecone-Kafka-Kubernetes stack a blog recommended." What do you build?
20M × 3 KB = 60 GB of raw vectors — one large-memory machine holds it uncompressed. Commit: single-node HNSW (no PQ needed — skip the recall loss you don't have to pay), one replica for failover, batch ingest nightly plus a small upsert queue, hybrid BM25 via the Postgres/Elastic they already run. No sharding, no Kafka, no dedicated cluster. Name the escalation trigger like a senior: "when the corpus passes ~100M chunks or QPS makes replica fan-out insufficient, we shard — and the eval suite we build now is what makes that migration safe." Simplicity plus a named trigger beats the blog stack.
What breaks: a support-agent tenant where 95% of documents are restricted to team leads. Agents report "search returns nothing" for topics they can partially see. Diagnose and fix.
Classic post-filter starvation: the pipeline retrieves top-K by similarity then applies ACLs, and since restricted docs dominate the corpus (and the topic), all K candidates get filtered to zero — while readable results sit just below the cutoff, never fetched. Confirm by logging pre-filter vs post-filter counts (100 → 0 is the smoking gun). Fix: push the ACL predicate into the index so the search skips invisible nodes but keeps traversing until it accumulates K visible results. Interim mitigation while that ships: over-fetch (top-1000) before filtering — but say why it's a bandage: it burns latency and still starves in the worst cases.
Requirement flip: freshness tightens from 5 minutes to 5 seconds (you're now indexing live chat and incident channels). What changes in the design, and what stays?
The index survives: HNSW upserts are already immediate — that's why we chose it. The enemy is the pipeline in front: batch embedding windows and queue latency eat the whole budget. Changes: stream per-message embedding with tiny batches (trade GPU efficiency for latency — say the price), a priority ingest lane for live sources so bulk backfills can't queue ahead of them, and a Lucene-style near-real-time pattern if needed: new vectors land in a small fresh segment searched brute-force (thousands of vectors — exact scan is free) and merge into the main index continuously. Freshness lag becomes a paged SLO, not a dashboard curiosity. What stays: hybrid retrieval, rerank, ACL pushdown, the eval loop — the quality machinery is freshness-agnostic.
A country estate outside Washington, and a task nobody was meant to finish.
You are handed a pile of parts, two helpers, and ten minutes to build a five-foot cube with seven-foot diagonals. The helpers are junior staff whose job is to become "increasingly lazy, recalcitrant, and insulting."
This was the Construction Test at Station S, where the Office of Strategic Services screened people for work behind enemy lines. Murray and MacKinnon's 1946 write-up gives it one flat sentence: "No candidate ever completed this task."
Not one. The cube was never the deliverable, because nobody could build it in the time. That is the design. Take the artifact off the table and the only thing left to record is behaviour: the brief is vague, the clock is short, the room is against you.
It was not cheap: fifty-plus professional staff at Station S, about 5,500 men and women through the centres, roughly twenty months of a war. The programme is the ancestor of the corporate assessment centre — and of your loop. Yours is buildable, and finishing counts; the vague brief and the mid-task shove carried over intact. Read the scoring before the subject.
The candidate records were ordered destroyed at the end of the war; a 2015 reanalysis of the one surviving matrix names the hole — the staff "did not know precisely what they were asked to predict."
Your interviewer says "now 100x the users" and your design falls over. What are they writing down? Write it before you read on.
The fraud desk was the product.
PayPal had the detector. Max Levchin's team built "PayPal's proprietary IGOR system". It worked, and it still needed a desk: a team of 100 employees, about one-sixth of the company, sat reading Igor's graphics, because "humans review the evidence and decide whether to freeze" an account. Igor flagged. A person decided.
A KDD paper later wrote the rule. Against an adversary, "the performance of a classifier can degrade rapidly after it is deployed", and the only remedy was "repeated, manual, ad hoc reconstruction". The un-automatable step was never solving the problem. It was finding the one you now had.
Palantir, founded in 2003 by Thiel and others, sells that step — an engineer posted beside the user to learn "how and why they make decisions" — which is why its interview hands you an infection and no architecture.
Read all four loops below the same way. Each is an instrument for the step its company cannot buy — so ask what each one still pays a human to do, instead of memorising four rubrics. Whatever a company automated, it stopped grading you on years ago.
Name the company you will actually interview at, and the one step in its work no tool it owns can do. Write it before you read on.
Bob Braden of USC's Information Sciences Institute did the arithmetic in November 1992. One question, one answer: at best a client waits "RTT + SPT" — one round trip plus server processing time. TCP charged "2*RTT + SPT, larger than the minimum by one RTT", because the three-way handshake must complete first. SPT is the only term the server's dashboards show; the extra trip is invisible from the machine that is late.
His target was not throughput. At least ten packets for two was "a high relative overhead", but "a more important consideration is the transaction latency": faster CPUs and fatter pipes shrink SPT; neither touches the RTT. "As CPU and network speeds increase, the relative significance of this extra transaction latency also increases."
A cold HTTPS request pays four such trips: find the address, open the line, prove who you are, then finally ask. This was the fight over one of them. It cost thirty-eight pages — RFC 1644, July 1994 — and it lost: that trip is the server's only proof the asking address is real, and T/TCP replaced it with a counter you could guess. Google bought it back with a cookie in 2011: a projected 15% off HTTP transaction network latency.
In those nineteen years the round trip never got cheaper. Bandwidth did. Only the count and the distance were ever negotiable.
Your dashboard says milliseconds; your user half a planet away waited seconds. Write down where the difference went.
Two Dorados, one wire, and a call that may arrive twice.
Andrew Birrell and Bruce Nelson's Cedar RPC package retransmits until acknowledged, so the callee sees the same call twice. They put the recogniser in the packet. The caller mints a call identifier — "the calling machine identifier …, a machine-relative identifier of the calling process, and a sequence number" — and the callee keeps "a table giving the sequence number of the last call invoked by each calling activity". Not greater, discarded. The Idempotency-Key header's mechanism, 1984, built for a flaky wire and not for money.
The mechanism was ten years old and Sun's NFS was hanging anyway. Mount a filesystem soft and a lost reply surfaced as a status programs ignored, so "system administrators tried to tune various parameters (time-out length, number of retries)". Waldo, Wyant, Wollrath and Kendall recorded it in 1994, in three words: "These efforts failed." Mount it hard instead and nothing failed, it waited — "one server crashes, and many workstations—even those apparently having nothing to do with that server—freeze."
They were not short of the fix. They were short of a field to put it in. The report's example is a queue, not a filesystem: "The only way to fix this is to change the protocol to add request IDs. But since this is a standardized interface, there is no way to do this."
Pick an endpoint you have shipped. If a client sends it twice, what is true after? Write it down first.
Two questions decide whether an index gets used; only one has an answer. Selinger's team wrote that one down in 1979 as a definition: predicates match an index when the columns they name are "an initial substring of the set of columns of the index key". Nothing is measured.
The second is how many rows the predicate returns. Nobody could know, so the paper guessed. It assumed "an even distribution of tuples among the index key values", multiplied predicates as if "column values are independent", and for an equality on an unindexed column wrote a flat 1/10.
Those are not stale numbers. They are the model. Run ANALYZE and the numbers get fresher, not different in kind: a value too rare to land in the sample is still handed the average share, and two columns that always move together are still multiplied as if independent. The plan moves with the estimate — your index skipped on a table whose schema never changed.
Those guesses shipped. Selinger's 2003 interviewer credited the paper with the principles that "underlie the query optimizers in every major database system today" — assumptions included. The authors knew their costs came out "often not accurate in absolute value". You never choose the plan. You choose the tables and the index column order; the optimizer chooses the rest from those guesses.
Take the hottest query in a system you have shipped. Write the index you would add, columns in order, before the chapter gives you the rule.
A deadline nobody could move, and a number nobody had computed.
The law required insurance marketplaces by January 1, 2014; federal enrollment opened 1 October. Two pre-launch assessments flagged capacity planning. A later audit: "Capacity requirements for hardware for the FFM system were not developed." Nobody had turned expected people into machines.
On 26 September, CMS officials tested at CGI Federal and found the site could carry "far fewer concurrent (simultaneous) users than planned". They drove to the hosting contractor and asked for double. A Terremark manager, later: "We were pulling gear from all over the world, renting planes to get hardware here that was intended for other clients."
Double of what, though. Every visitor entered through one account gate, the EIDM. The contractor who built it "did not know that all visitors to the website would have to enter through the EIDM system and, therefore, underestimated the capacity needed." The gate took the whole arrival rate. It had been sized as though it took a slice.
CMS launched "without the required verification that it met performance requirements." Day one: 250,000 concurrent users, by the U.S. CTO's count five times the number anticipated. Outages began within two hours. Six people finished an application. By March 2014, CMS had committed $840 million.
They could move hardware across the planet in seventy-two hours. What they could not produce was the one number to multiply by two.
Which number should they have computed, and what would it have changed?
One big machine, or many? Two papers, one AFIPS volume. Daniel Slotnick, later an ILLIAC IV co-author, at pages 477–481: "Unconventional systems." Gene Amdahl at 483–485, IBM in Sunnyvale, promising to demonstrate "the weaknesses of the multiple processor approach". One figure, no equations.
Housekeeping is sequential, he argued: 40% of executed instructions in production runs, halved at best. So: "Overhead alone would then place an upper limit on throughput of five to seven times the sequential processing rate, even if the housekeeping were done in a separate processor." The processor count does not appear. It cannot.
Then he priced two processors sharing memory: 2.2 times the hardware, two tenths of it their shared crossbar. After memory conflicts, about 1.8 — "a price performance degradation to 0.8 rather than an improvement to 2.0 for the single larger processor."
ILLIAC IV went ahead anyway. Planned: 256 processors under central control, four quadrants of 64, a billion instructions a second. The April 1972 report: "cost escalation and schedule delays" had left one quadrant, 64 processors, about 200 million instructions a second. One control unit, one instruction stream, every ALU lock-stepped: "only those PEs in an enabled state are able to execute the current instruction." A data-dependent branch parked the rest; each processor earns only where all agree.
Nobody has engineered around it. The only lever is how much of the work must agree.
One server at its limit, budget for nine more. Write down what in your system has to agree.
Both main clearing banks for US government securities operated a few blocks from the World Trade Center. Bank of New York's funds transfer and bond settlement systems sat at 101 Barclay Street, one block north. Both its buildings were evacuated.
11 September 2001: Roger Ferguson, the Fed's vice chairman, was the only Board governor in Washington. The Board's whole statement was two sentences: "The Federal Reserve System is open and operating. The discount window is available to meet liquidity needs." Within the week the bank was reported overdue on $100 billion in payments. The Fed injected more than $100 billion in additional liquidity.
Then regulators came for the back-up sites: live copies of the ledger, fed as payments land. A 2002 draft asked whether to require 200-300 miles of separation. The final paper, 7 April 2003, dropped the mileage — the agencies would not "prescribe specific mileage requirements" — and set a clock: two hours, inside the business day.
One clause did the damage. Core clearing organisations' copies must sit "well outside of the current synchronous range". Three federal agencies could name no distance that both survives a citywide disaster and lets a commit wait for the far copy. So they put the copy beyond the reach of waiting, and left the organisations to measure — and defend to an examiner — how far behind it runs.
Your far copy is deliberately too far to wait for. Write down what that costs you the second the primary dies.
A database nobody could afford to re-cut.
Google's advertising backend ran on MySQL "manually sharded many ways". One column decided everything: the scheme "assigned each customer and all related data to a fixed shard". Every campaign and ad group an advertiser owned sat together, so per-customer indexes and joins stayed cheap — and any question that was not about one customer needed "knowledge of the sharding in application business logic".
Replication was not room. The team’s own limitations slide filed master–slave under availability, "downtime during failover"; under scaling it wrote "Grow by adding shards".
Adding shards meant re-deciding which shard each customer belonged to and moving them live, on the database that bills the advertisers. The Spanner paper puts the invoice on the record: the last resharding "took over two years of intense effort, and involved coordination and testing across dozens of teams to minimize risk."
So they stopped paying it. "This operation was too complex to do regularly" — the team capped growth on MySQL and pushed the overflow into external Bigtables, which "compromised transactional behavior and the ability to query across all data." The dataset was tens of terabytes: not too big for the hardware, only for the cut. One column, chosen once, had become a ceiling on how much of Google's ad business its own database was allowed to hold.
Two years to move the rows once. What would you build on day one so the second move is cheap? Write it before you read on.
It began as a calendar.
David Gifford was a graduate student at Stanford University building "a distributed calendar system" called Violet, keeping each file on several PARC machines so one dead server did not take the calendar down. A rule for that already existed, and it worked: update a majority of the copies.
He broke it over latency. A majority is one number, and one number prices reads and writes identically, while a file's "read to write ratio" is almost never one. A system directory is read constantly and written almost never. The copy on your own disk answers sooner than any copy across the network. One count cannot say either thing. So Gifford stopped counting copies, gave each copy votes, and split the count in two: r votes to read, w votes to write, set separately, per file.
Cheap reads were now purchasable. His own table names the price: r=1, w=3 across three servers reads a directory in 75 milliseconds from one machine; every write costs 750 milliseconds and dies if a single one of the three blinks.
Which numbers, then, for which data? CAP, consistency levels, quorums, conflict resolution — everything below is machinery for a question the arithmetic cannot answer, because a directory and a shopping cart want different numbers. Gifford specified the guarantee and left the numbers to "knowledgeable system administrators", naming no method.
Majority-of-copies is one number, and it worked. What can r and w say separately that one number cannot say at all?
The consumers were ambulances. London has only so many ambulances.
The London Ambulance Service switched its Computer Aided Despatch system fully on at 07:00 hours 26 October. Calls became incidents, incidents went on a list, and when a vehicle's status came back wrong the system raised an exception message and listed that too.
The inquiry found those days "were not exceptionally busy days in terms of emergency incidents or patients carried". Nothing crashed: "the computer system itself did not fail in a technical sense". What grew was the buffer. "As the exception message queue grew the system slowed" — unrectified messages generated more messages, and messages scrolled off the top of the operators' screens. Ambulances arrived late, or twice. Callers rang back, which "greatly increased the number of call backs": the queue was making its own producers. Average ring times peaked at around 10 minutes.
So they drained it: "the exception report queue was cleared in an effort to increase the speed of the system" — and uncovered incidents probably went up, because clearing the queue did not clear the incidents in it. By the afternoon of 27 October London was dispatching off printouts again. The £937,463 winning bid was never the bottleneck. The consumer at the end of that queue was an ambulance crossing London, and nothing you can write makes it arrive sooner.
The servers were healthy and the day was quiet. Which one number would you have put on the wall, and at what value should it page?
The books balanced. The books were wrong.
A subpostmistress scanned one pouch of £8,000 out of her core branch into her outreach branch. Horizon recorded the transfer four times.
The retry was honest. Her log-on script had not closed down, so the pouch delivery script "thought it had not finished and attempted to repeat the last part of the script" — each repeat wrote a fresh remittance, with nothing recording that the first had been written. At Callendar Square in 2005 the same missing memory ran backwards: transactions vanished off the counter, staff re-keyed them, transfers doubled.
Nothing objected. Each phantom remittance was recorded exactly like the real one, so the accounts agreed with themselves all the way down. Fujitsu's chief architect told the High Court the error was "not subject to receipts and payments problems" because the bug "had the effect of making it look as though a user was simply doing something multiple times".
Only one instrument could see it, and it was not in the software. Her outreach branch held Horizon receipts for £32,000 against the £8,000 in cash she actually carried: £24,000 short. Fujitsu then found 112 occurrences across 88 branches in five years.
No internal invariant fired, because none could: a ledger can only check itself against itself. Internal agreement is not evidence. Until someone counts the drawer, the books are a record of what the software believed.
Every entry balanced and the retry was honest. What should the receiver have remembered, and what could catch it once it didn't?
A spare for everything.
AT&T's network was redundant at two radii: inside each switch, the Direct Link Node onto the signalling network was duplicated; across the country, 114 4ESS switches. At approximately 2:30 p.m. EST on Monday, January 15, New York hit a minor hardware problem and did what it was built to do: suspended calls, ran recovery. Four to six seconds. Invisible.
Then it came back. Its first call attempt told a neighbour's DLN New York was up. Mid-rewrite of its status map, two more landed within an interval of 1/100th of a second. Data damaged; the processor dropped out of service. Its duplicate mate took over, caught the next pair, died the same death. That neighbour's return did it to everyone connected to it — "a chain reaction", said AT&T's Larry Seese.
Redundancy inside a switch and redundancy across a continent died the same death. Both sides of both walls ran the identical mid-December program — an update built to reach the backup signalling network faster — and the trigger crossed on the links every switch shared.
Standard recovery "proved inadequate": each returns a switch to service, and coming back was the weapon. What worked was subtraction: suspend signalling on the backup links, starve the cascade. Last link cleared 11:30 p.m. EST. Once the walls are drawn, the only lever left is making less cross them.
Name the thing that is identical on both sides of every wall you drew. Write it down first.
EVE Online ships the partition key every architect wants.
One universe cut into over 7,500 solar systems, no hash function anywhere; CCP reinforces hardware one solar system at a time. It still burned.
On 27 January 2014 two coalitions met in B-R5RB, Immensea: 6,058 pilots inside one key.
Cache it? A simulation has no cacheable read side: every read is somebody's next tick, and a tick is a write. Replicate it? Every replica applies every tick: work multiplied, not divided. Split it? Everything in a solar system must see everything else, in order; cut it across two machines and you get no fight at all.
The precondition nobody states: splitting a key into N works on a like counter because the parts sum and order does not matter; it is illegal the moment the parts must interact or be ordered. B-R5RB is the far write side of that fork — all writes, nothing to absorb, no legal cut.
Only the floor was left: degrade the key on purpose. CCP Veritas announced Time Dilation in 2011 with the ceiling up front — "Time dilation does not solve everything." B-R5RB ran at 10% speed for 21 hours, executing "just over 2 hours worth of simulation"; CCP put the wreck at $300,000-$330,000. The clock was the only thing left to spend.
The chapter is about to split a hot key into N pieces. What must be true of what lives under that key for the split to be legal? Name one key where it is not.
A switch that cannot run out of buffers is cheap. John Nagle priced one: four 56Kb links, a 15-second lifetime ceiling, 420K bytes — so it “need never drop a packet due to buffer exhaustion”.
RFC 970 proves that is the doomed switch. The queue grows until transit through it exceeds the packets' time-to-live. What gets forwarded leaves with a lifetime of “exactly one”; the next hop zeroes it and discards. Infinite storage, first-in-first-out, a finite lifetime: the network “will, under overload, drop all packets” — and “not an artifact of the infinite-buffer assumption”.
Time runs out, not space: memory bought only distance between arrival and answer. Nagle queued 916 datagrams and stopped there, his connection “timed out due to excessive round trip time”. The caller left while the buffer was still winning. A bound is not a concession; it is the only depth at which a queue helps.
The switch was not mute: Nagle's remedy has it warn a host with a Source Quench, “while minimal, is sufficient”. But an advisory is not a refusal, and it lost: deprecated for routers in 1995, transports in May 2012, “an ineffective (and unfair) antidote for congestion”. What survives is loss, indistinguishable from a broken wire. You hold the connection: your refusal gets a status code, a Retry-After, a name for who loses first — all chosen before the queue grows.
Your only refusal is dropping the packet. Name the channel your service has instead, and what it buys.
A wall of working instruments, telling four men the wrong thing.
At 4:00 a.m. on March 28, 1979, feedwater to TMI-2 stopped. A pressurizer relief valve opened as designed. Thirteen seconds in, it was commanded shut. It didn't.
They had a lamp for that valve, wired to the solenoid that carried the command, not to the valve stem. The Commission: the light "showed only that the signal had been sent to close the PORV rather than the fact that the PORV remained open."
The board read SHUT for 2 hours and 22 minutes while the plant bled out: 32,000 gallons, over a third of the coolant, in the first 100 minutes. The Commission priced it at $1 to $2 billion, all of it behind a lamp that read shut.
More signal did not help: a shift supervisor testified "there had never been less than 52 alarms lit in the control room" on an ordinary day, and Craig Faust wanted the panel thrown away.
The men were not under-instrumented. They read a working instrument correctly — one tapped a step short of the effect it stood for. It reported a command leaving the room; the reactor was on the other side of the wire.
A lamp is only as true as the point it is tapped. Ask where your checkout-success number is tapped — everything that fails past it reads green.
Which of your alerts measures a command you sent, not an effect a user got? Name it first.
P-waves arrive first and do little harm. S-waves bring the ceiling down. The gap between them is the whole warning.
Japan's early warning system cannot choose when to stop listening. The S-wave sets the deadline, a consumer whose clock the pipeline does not control, so the system declares a cutoff and sizes the earthquake from what it has. A pick is a station, a time, an amplitude; nothing on it says which rupture produced it. That must be inferred.
In January 2018 two earthquakes broke at the same moment, about 400 km apart. IPF, the associator JMA added for this case, weighing amplitudes as well as travel-time differences, located both correctly. The step after it did not: the M4.5's magnitude was calculated with displacement amplitudes of the M4.0, and the predicted shaking came out at least two intensity units too strong.
Nothing was late or lost. Every pick was a real arrival, correctly timed, filed against the wrong earthquake: a reading that belonged to a population the clock had already closed. The cutoff was per stream; completeness is per source. One watermark over a merged input is the same defect, one clock speaking for two populations, which is why the declaration belongs per partition and why a silent partition holds everyone's windows open. Slack cannot repair it: waiting longer only lets more of the other earthquake in.
Your stream merges mobile, web and a partner's nightly upload. Where do you put the cutoff, and what happens to an event that lands after it?
The write-up, that year, of work already shipped.
Doug McIlroy’s problem at Bell Labs was arithmetic. Unix spell’s word list ran to about 250,000 bytes of clear text. A PDP-11 could address 32,768 sixteen-bit words.
So he stopped storing words. Dennis Ritchie supplied “a simple superimposed code scheme first proposed by Bloom”: 25,000 words across 400,000 bits, eleven hash functions, false acceptance under 1 in 2,000. Sixteen bits a word.
Then “experience showed, however, that 25,000 words wasn’t enough”. Thirty thousand wanted a bigger table and there was none. The rate is words over bits; only one was his to move.
So he priced the other option — not the words, their hashes. Twenty-seven bits apiece, for an error of 1 in 4,096; sorted, stored as differences under a variable-length code. 13.60 bits a word, 14.01 once split across 512 bins for fast lookup — “a tight, but workable fit.”
More words than the filter held, a lower error rate than it gave, at fewer bits a word. A filter never competes with storing the items. It competes with a well-coded index of the fingerprints you would keep instead: 13.57 bits an entry, against the 17.5 a filter needs at the same error. Tighten the error; the gap widens. The twenty-to-one on the next page is against a hash set, not this.
Price the exact alternative to your filter — not a hash set, a compressed index of its hashes. At your error target, which is actually smaller?
A bank that took the maintenance window to its maximum and still had no way back.
TSB had been divested from Lloyds Banking Group and still rented Lloyds' core banking system. Its Spanish owner, Sabadell, wanted it on Proteo4UK, by its usual method: all the data, one weekend.
So the window went to its maximum. Not four hours at 2 a.m. but the whole dress-rehearsed weekend of 20–22 April 2018, 5.2 million customers, reader and writer changing in the same instant. No stretch of writing both stores, no week counting how often they disagreed. In week one 20–30% of retail and business customers could not pay anyone online; the regulators' fine was £48,650,000.
But the way back shut before a row moved. To migrate at all, TSB had to notify Lloyds it was exiting, waiving its right to a copy of Lloyds' platform, good until 31 March 2019. TSB's board gave that up on 10 April 2018, eight days before it voted whether to go at all; after the data moved, "a very short window" remained before the Lloyds platform was disabled. The FCA's word is "irreversible"; a roll-back plan "was not possible".
Every rollback in the four phases ahead — flag off, truncate, flip the reads back, keep writing the old store for a fortnight — is a return to the old store. TSB no longer had one.
Name one thing TSB could have shipped first that would have left a way back. Write it before you read on.
No network, no disk, one keypress of budget.
Dale Grover, Martin King and Clifford Kushler are making a nine-key pad type English. Press 2-2-7-3 and you have asked for a word, not a spelling. Something must rank the candidates, and the device cannot: no server, no memory to spare, the answer due before your thumb lifts.
So the ranking is finished years before the key is pressed, and ships as storage order. Their patent, granted 6 October 1998, is blunt: "Within the memory, the words are stored in order of decreasing frequency of use so that the proper order is automatically presented." Nothing scores at keypress time, because there is no time at keypress time.
Look at the price. Memory being "a critical factor for mobile phones", the stored lookup tags are dropped and regenerated during the search, so the read walks every word of the right length. A scan on the hottest path, chosen on purpose: memory was scarcer than ranking, which was already free. Same kernel, inverted cost curve: this chapter spends memory to buy a read that never scans.
Now count the traffic that forced it. One user. One device. One thumb. Zero requests per second, and the order still had to be settled before the key went down. Volume never entered it. The one thing that could not stretch was the budget inside a single request, one keypress wide.
What forces precompute — the number of requests, or the budget inside one? Write it before you read on.
A cardiac monitor is a notification system with one recipient.
A hospital ward can raise several hundred alarm signals per patient per day, and an estimated 85 to 99 percent of them need no clinical action. On another hospital's 15-bed unit, staff documented an average of 942 alarms a day — about one critical alarm every 90 seconds. Correct or not, every one spends the same finite thing: the attention of whoever is on the floor.
So the floor defends itself. Clinicians turn the volume down, switch alarms off, widen the limits past safe. The Joint Commission's 2013 alarm-safety alert counted 98 alarm-related events in its own database, 80 of them deaths; the contributing factor it counted most often was alarm signals inappropriately turned off.
In January 2010 an 89-year-old man's heart rate fell for 20 minutes and stopped. Lower alarms beeped at the station and crossed three hallway signs. The ten nurses on duty recalled none of it; staff on the unit told investigators they were desensitised. The crisis alarm, the one signal built to cut through, was off. Investigators never said why. The family settled for $850,000.
The channel was never saturated; the receiver was, and the routine traffic had spent it first. Ration the routine traffic or lose the critical one — the last hop is a person you do not own and cannot scale.
Your bulk lane and your critical lane share nothing but the person holding the phone. What has bulk already spent?
Staff at the CSNET Coordination and Information Center kept fielding one question: why is a single mail message delivered several times? Craig Partridge traced it: no mailer was broken, the hole was in SMTP.
A sender ends the message with a single dot; the receiver replies 250. In between, says his February 1988 memo RFC 1047, "the message is active at both the sending and receiving mailer" — the sender must assume it was not delivered, the receiver must assume it was. Break the link there and both copies go out.
Half his fix is this chapter's rule: reply as soon as the message is "safely put in a non-volatile (e.g., disk) queue." Persist, then ack. The other half he answered with a clock rather than a name. Shrink the gap; let the sender wait longer before giving up — ten minutes, he wrote, would not be unreasonable. RFC 1123 made exactly that a floor in October 1989; RFC 5321 still carries it, and cites him, in 2008.
Nobody ever wrote the other rule: an identifier on the message a receiver is obliged to remember. The standard's whole remedy is still to shrink the window, and a window you only shrink is a window. Mail still arrives twice.
"Design WhatsApp," says the interviewer. Your chat server has the same window, and a longer timeout only narrows it. What must be true before the single tick — and what must the phone send so a retry is not a second message?
By 2017 Weibo had bought its way out of capacity. With Alibaba Cloud it built DCP, a container platform with one target: 1,000 nodes in ten minutes. Exercised at the peak that sits on a calendar — the 2016 Spring Festival, more than 1,000 ECS instances created hours beforehand. Whatever came, the fleet could grow into it.
Then noon, Sunday 8 October 2017. Lu Han posted one sentence: "Hi everyone, I want to introduce my girlfriend @GuanXiaoTong to you." What saturated was not delivering it but everyone converging on it: so many fans commented that the platform was inaccessible for two hours.
Both directions of that pipe are the same shape. Traffic in, or one pushed copy per follower out — the job's size is set by a number in someone else's account; elasticity only spreads it over more machines. Ten minutes to a thousand machines makes the work arrive faster. It never makes it smaller. Which leaves one lever: decline to do it.
And the follower count you decline above is a reading of your own graph. A 2012 study of 80.8 million Weibo profiles and 7.2 billion follow relations found following preference more concentrated and hierarchical than Twitter's — more of the graph aimed at fewer accounts. Borrow a threshold from one tail and you've sized for the wrong one.
Forty-five minutes, a whiteboard, one tap that becomes a hundred million writes. Where would you refuse to do the work at all? Write it before you read on.
The store was the easy half then, too.
Ray Ozzie's document store worked. Maybe six months after the team formed in January ’85, the hard part was elsewhere: they "needed to take the concept of synchronization and replication much more seriously", Ozzie told the Computer History Museum in 2020, so they hired Steve Beckhardt for it. Replication was not a feature of the store. It was a second system.
It had to be. There was no internet to lean on: Notes servers dialled each other and traded documents store and forward, so replicas drifted unseen between calls. Two people editing one document apart was not the edge case the design had to survive. It was the base case.
Notes does not guess. A conflict is two replicas that each saved the document between replications — what the two writers could see, not when they typed. $Revisions, "which tracks the date and time of each document editing session", only names a winner; the loser survives untouched, titled "Replication or Save Conflict".
Notes can read its documents’ fields; a file sync tool cannot. The Domino help offers to merge them, "if no fields conflict". Touch the same field twice and it withdraws. Structure buys you every field that did not collide. On the one that did, it buys nothing.
Design Dropbox, says the interviewer. Two devices edit one file offline, neither edit may be lost. Where does the losing copy go, and who finds it? Write it down first.
Indian Railways opens its Tatkal quota one day ahead, on the stroke of 10:00 for AC berths. In August 2016 a developer published a script that reads the Date header off IRCTC's login page and sets your clock to the server's “as close as ±1 second” — so your claim fires inside the opening second.
A Tatkal claim is a hold, and the berth stays blocked while IRCTC hands you to a bank's gateway it does not own, cannot hurry, cannot roll back. When the hold dies before the bank answers, IRCTC's advisory names the result: “money debited from the account but ticket not booked”. No undo, only compensation — IRCTC returns the amount “on next day”, the bank credits the passenger “within 2-3 working days”.
Railways widened the pipe: connections 40,000 to 300,000, capacity 2,000 to 15,000 tickets a minute. The race did not settle. On 19 January 2016 the Ministry conceded scripting “cannot be stopped” and answered by slowing everybody down: a “minimum payment time check” sitting inside the hold. Their own floor on a booking: 35 seconds. Friction aimed at the bot was time the honest buyer's berth spent blocked.
So the timer has no right value. Cut it and a buyer is debited for a berth he never gets, then waits days for his money back. Stretch it and the berth sits locked by whoever reached it first.
Your hold expires and the gateway has not answered. Write what your system does.
The web agreed on a fence, never a speed limit.
The robots.txt consensus of 30 June 1994 shipped two fields: User-agent and Disallow. Unrecognised lines, ignored. Its status section: "It is not enforced by anybody."
Twenty-eight years later the IETF ratified it as RFC 9309, Standards Track, September 2022. It gained Allow and a rule that a crawler "SHOULD NOT use the cached version for more than 24 hours" — your robots.txt TTL. The strings "crawl-delay" and "rate" occur nowhere in it.
A server can say where you may not go. It has never had a way to say how fast. Most hosts therefore say nothing about speed: the format gave them no field.
So the rate is not a value you read. You compute it, and it moves: when you last touched this host, how slowly it answered, whether it pushed back. Per-host state, rewritten by the very fetch it governs. The base term in your max() is there for the silent majority.
Nobody on the far end audits one request against a declared rate; none was declared. Enforcement arrives late and blunt: a ban, not a warning. So the clock is yours to keep, one per host — inside the same queue that must also decide which host is worth reaching first.
One structure, two jobs: a per-host clock and a ranking of what matters. Sketch a layout that cannot break the clock even when the ranking screams. Write it before you read on.
A bank's day ends when the overnight batch finishes, not when the branches close.
One mainframe scheduler ran that batch for NatWest and Ulster Bank. It queued each job and ensured "each job is processed in the correct sequence." Nothing in it ever had to decide which machine fired a due job: there was only one.
On Tuesday 19 June Technology Services "backed out the software upgrade" on that scheduler — and "a significant number of jobs failed to appear in the batch queues".
There was no replay button. Support staff "focused on manually re-loading jobs into the batch queues": by hand, in sequence, because out of order gives wrong balances and wrong interest — while every midnight stacked a fresh batch on the one they had not cleared.
6.5 million customers found out what a hand-replayed batch leaves behind. The banks "had not processed their standing orders on time", had applied "incorrect credit and debit interest", and had left "duplicated entries on their statements". £70.3 million went back in redress.
One statement could carry both halves: a payment that never went out, and lower down the same payment made twice. A re-loaded job carried nothing that said it was not the job that had already run.
You are the candidate: "design a job scheduler — cron, but distributed. Millions of jobs, a fleet of machines." Zero fires and two fires just landed in one account. Which should your design make impossible, and which merely survivable? Write it first.
Approximation was the easy half; the error's sign was the argument.
AT&T's backbone, sampled and aggregated, ran to 500 Gbytes a day, about ten billion fifty-byte records. Gigascope read such streams live, peaking at 1.2 million packets per second. A counter per flow is not expensive there; it is unavailable. Fixed array, one pass.
Giving up exactness was nobody's insight; earlier sketches had done that. What none had fixed was direction — their estimates could land below the truth as readily as above. For a count that is inaccuracy; for a ranking it is deletion. An undercount does not make a winner's number wrong, it removes the winner, and the list carries no mark where one fell out.
In June 2003 Graham Cormode and S. Muthukrishnan, the latter listed at "Rutgers University and AT&T Research", filed a DIMACS tech report that chose the sign on purpose: the bound is "one-sided, as opposed to all previous constructions which gave two-sided errors." One word buys a sentence no earlier sketch could write: "absolute certainty that this procedure will find and output every heavy hitter".
The sign is not free: because the error only adds, there is nothing to subtract back. Every key carries a floor bounded by ε‖a‖₁ with probability 1−δ — this chapter's (e/w)×N, traffic crossed, not keys counted. A worm raises that floor under every flow you built the thing to watch.
Top ten hashtags, last hour, live. What is one-sided error worth at rank 10, and at rank 1000?
The count is the invoice. Ask who kept the events.
Nielsen's ratings price American television advertising. They are not a count. They are a panel: about 40,000 metered US homes standing in for everyone else.
The pandemic severely curtailed Nielsen's in-person visits to those homes. Homes that tripped alerts stayed in the sample instead of being pulled. By February 2021 about 9,400 of those homes carried alerts that would normally have had them withheld.
So the Media Rating Council, the industry's measurement auditor, asked for the real number. There wasn't one. The missed viewing was never recorded. All Nielsen could do was re-run February without the suspect homes, and MRC published the gap: viewing among adults 18–49 understated by "approximately 2 to 6 percent". A simulation carrying standard errors. Not a recount.
On 1 September 2021 the MRC board voted to suspend Nielsen's national and local TV accreditation. "While we are disappointed that the situation has come to this, we believe these are the proper actions for the MRC to take at this time," said George W. Ivie. In March 2022 Nielsen agreed to go private for $16 billion; national accreditation returned 17 April 2023, local did not.
In January 2022 a trade body turned those percentages into lost impressions; MRC issued a correction: it "did not position these figures as adjustment factors". Nobody could ever say what the right number had been.
An advertiser disputes a bill from a year ago. Write down what you must have kept to settle it.
Fedwire settles government bonds delivery-versus-payment, so every delivery into a bank's account automatically debits that bank's reserve account to pay the seller. The wire commits the effect; the bank commits the record. On Thursday, 21 November 1985 the Bank of New York found its database of accounts and transactions damaged and could not tell the Fed where to forward the bonds arriving for its customers. They kept arriving. The effect committed. The record did not.
By evening the overdraft was $32 billion against total assets of $15.9 billion; overnight the Fed lent $23.6 billion, against every asset it owned and its customers' securities. Interest: about $5 million, near 7 percent of nine months' earnings.
None of it was missing. It was unrecorded — yet every movement was already recorded, in order, never overwritten, in the reserve account the Fed kept. The counterparty's ledger held the money's truth throughout. The bank had lost not the money but anything of its own to match against it; and a bank that cannot read its own record cannot say whether a seller was paid once or twice.
They rebuilt Wednesday's trades by 5 a.m., Thursday's by 10; Friday's were already stacking in the same account. At about 11:30 a.m. Corrigan stopped the wire into the bank, because nothing else could stop the record falling further behind the money.
You are about to design a payment system. Before you read on: name the two things its record must never get wrong, and who checks it.
The cache had a repair system. The repair ran through the database.
An automated system watched the cache for invalid configuration values and refetched them from the persistent store. Robert Johnson's postmortem names the assumption: it "works well for a transient problem with the cache, but it doesn't work when the persistent store is invalid."
On 23 September 2010 a change made the stored copy of one value read as invalid. Now "every single client saw the invalid value and attempted to fix it" — and every fix was a database query. The cluster took "hundreds of thousands of queries a second."
Then the loop closed. A client whose query errored read that as one more invalid value and "deleted the corresponding cache key" — so it missed again, and queried again. Refilling was the load, and it fed itself: "even after the original problem had been fixed, the stream of queries continued."
A miss is supposed to cost one database read. It costs one read only while the database can still answer. Under that line a cold cache refills through the thing its own misses buried, so time stops helping; the outage outlives its cause. Facebook was dark about 2.5 hours. The only exit was removing the traffic: "we had to stop all traffic to this database cluster, which meant turning off the site."
You build the cache now — no Redis, no Memcached, clock running. What has to be true of the database behind you before a lost key is "just a miss"?
After a crash, the disk was not the database.
Pages reached disk whenever the buffer manager chose. Some held uncommitted changes. Others lacked changes from transactions that had committed and told the user so — commit forces the log, not the pages. Nothing on the page said which.
The obvious repair is selective redo: replay only what's missing. Mohan, Haderle, Lindsay, Pirahesh and Schwarz showed it cannot hold. Every page stores the log position of the last record applied to it: one log, thousands of pages, each at its own point. Skip a record and that number lies. "By not repeating history, the page-LSN is no longer a true indicator of the current state of the page." Undo then reverses an update the page never took.
ARIES, their answer, stopped being clever. From the last checkpoint, redo everything in order, including work it knows must be undone — "even the missing updates of the so-called loser transactions" — then undo. Reading a record never consumes it; restart is itself restartable. That "reestablishes the state of the database as of the time of the system failure." The store is demoted: a stale view of a log prefix that carries the number saying which prefix. By 1999 it had reached DB2, NTFS and MQSeries.
One thing no replay can rebuild: whatever an acknowledgement outran on its way to stable storage.
Your store is built by consuming a log; its disk copy is stale by an unknown amount. What must it write beside the data? Write it first.
At 14:14 EDT the alarm process in FirstEnergy's grid control system stalled — two processes writing one data structure at once. GE's Mike Unum: the corruption sent it "into an infinite loop and spinning". The chart printers kept scrolling, but "the chart pens showed only the last valid measurement recorded". A flat line reads as a steady one.
At 14:41 the primary server died under the backlog. The standby took over exactly as designed, the stalled alarm application intact. It died 13 minutes later. Thirteen minutes was the whole value of the second machine, because the fault was never in the machine — it was in the state the application carried, and the standby inherited the same corrupted structure and the same growing queue. Unwatched, the grid came apart: $4 to $10 billion in United States costs.
Redundancy answers one question: does this box fail independently of that one. When the fault is in what every copy holds, copies do not vote — they queue, and they die in order. Two replicas of a metrics shard hold the same set of series; the label that mints fifty million of them overruns one replica's memory and its partner's in the same minute, with the rebuild still replaying. Adding a replica adds a second victim.
You run two replicas of every shard. Name a failure where the second replica dies of the same cause as the first — and what redundancy would have to look like to survive it.
The cheap code already met the target. Repair cost it the argument.
Cheng Huang's team had the answer on the table. Azure held three full copies of everything not yet coded; Reed-Solomon (12, 4) — twelve data fragments, four parity — would settle sealed data at 1.33x, exactly the target.
They threw it out on a number nobody had been budgeting: "to do the reconstruction we would need to read from a set of 12 fragments." One fragment offline, twelve reads dragged across the fabric — and at fleet scale, something always is. Fan-in is not maintenance you schedule; it is traffic you pay continuously, on the customers' own network.
So they bought it down structurally. Split the twelve data fragments into two groups of six, one local parity per group, two global parities across all twelve. Sixteen fragments, still 1.33x — but one lost data fragment now rebuilds from six reads, not twelve. Halving the fragments read roughly halved the work: 893 ms to 418 ms on a 4 MB reconstruction. Repair is a bandwidth problem.
The bill is in the fine print. Local Reconstruction Codes are not maximum-distance-separable. Reed-Solomon survives any four losses; LRC survives any three, past that only the four-loss patterns that happen to be decodable. Same overhead, half the repair reads, and a shorter list of failures it is allowed to have. They picked which failures to lose.
Suppose you cannot trade fault tolerance away. What else makes a rebuild finish faster? Write it before you read on.
Before this, video came off a media server: Flash "streams video at a constant rate using the proprietary Flash Media Server". A session per viewer, held open for the whole show. You bought capacity by the concurrent stream, and no cache helped: no two viewers fetched the same object.
Move Networks sold video with no video server. It cut the file into what its patent calls "streamlets", "between about 0.1 and 5 seconds", encoded at every bitrate. Not a stream. A grid, inert: streamlets "may essentially be static files", so "no specialized media server or server-side intelligence is required", and "cached by cache servers of Internet Service Providers". Capacity now scales with distinct bytes, not viewers: the CDN bill.
Everything follows from who switches. The claim puts it plainly: "each of the end user stations initiate each of the shifts between the different bit rates". If the player picks per chunk, the encoder cannot know which rung it will want, so it must build them all first.
ABC shipped full-length HD on it, 24 July 2007; Apple republished the shape, 1 May 2009.
EchoStar bought Move's assets on the last day of 2010 for $45 million. The grid outlived the company. It still has to be built whole before anyone presses play, on a user-generated firehose where most videos never find one.
Every rung costs compute before the first play; most uploads are watched by nobody. Which would you build up front? Write it first.
Roger Weber and Hans-J. Schek at ETH Zurich, with Stephen Blott at Bell Laboratories, already had the VA-File, written up a year earlier. Then they measured an R*-tree and an X-tree against a plain sequential scan.
The trees lost "by orders of magnitude" on 50,000 images of 45-dimensional colour vectors. Fifty thousand: not a scale problem anyone outruns. Above around 10 dimensions the scan wins on average; any partitioning scheme must eventually "degenerate to a sequential scan through all their blocks", because almost every cell is empty and nothing is left to prune.
The year after, Beyer, Goldstein, Ramakrishnan and Shaft at Wisconsin showed what emptiness does to the answer: "the distance to the nearest data point approaches the distance to the farthest data point", on real data "for as few as 10-15 dimensions". Nearest stops distinguishing anything, and nothing says so. The search returns ten ranked rows with scores, exactly as when the ranking is real.
Which is how a decade of indexes went unchallenged. Most published techniques, they noted, "are not evaluated versus simple linear scan", and read carefully, their own reported experiments often show the scan winning. The failure sat in their printed numbers, looking like nothing. Only the dumb baseline run alongside exposes it, and a search that answers every query with ten ranked rows will never make you run it.
Your retrieval returns ten scored passages for every query, including ones whose answer is not in the corpus. How would you find out?