The gist
The requirements nobody writes down
A feature request says what an app should do: show a timeline, send a message, run a report. It rarely says how fast, how often it may break, how much load it must take or how easy it must be to change later. Those are nonfunctional requirements, and Chapter 2 is about making them concrete enough to design for.
The chapter runs one example all the way through: a social network's home timeline, where opening the app shows recent posts from everyone you follow. You can build it by querying at read time, which gets expensive for people who follow thousands of accounts. Or you can precompute each user's timeline whenever someone posts, called fan-out on write, which makes reading cheap and posting expensive, especially for an account with millions of followers. Neither is free, and deciding between them is exactly what the rest of the chapter gives you tools for.
Requirement 01
Performance: look at the tail
Two numbers describe performance. Throughput is how much work gets done per second. Response time is how long one request takes from the user's point of view, including network delays and time spent waiting in a queue. The chapter keeps these apart from latency, the time a request spends waiting before anything works on it.
Response time isn't one number, because the same request can take 40 ms one moment and a second the next. So the chapter argues for percentiles instead of averages. The median (p50) is what a typical user sees. The p95, p99 and p999 are what your slowest requests look like, and those are often your heaviest users, the ones with the most data.
Why the tail gets worse
Two effects make slow requests matter more than their share suggests:
- Tail latency amplification. If one page needs answers from ten backend calls, it's as slow as the slowest of the ten. A 1-in-100 slow call turns into a much more common slow page.
- Queueing. A few slow requests can hold up the fast ones behind them, so measuring on the server alone can make things look better than users experience.
That's why service agreements are written in percentiles, like "p99 under 1 second", rather than averages.
My take
The async ETL pipeline I built was a throughput story, not a response-time one. It didn't make any single API call faster. It kept many calls in flight at once, so more wallets got processed per minute. Having separate words for the two made it easier to say what actually improved.
My take
My AWS Lambda monitoring reported whether each crawler run passed or failed. After this chapter I'd also track how long runs take, as percentiles. A run that succeeds but takes three times longer than usual is an early warning, and a pass/fail check can't see it.
Requirement 02
Reliability: faults aren't failures
The chapter draws a line I'll keep using: a fault is one part going wrong, like a disk, a process or a network link. A failure is the system as a whole no longer giving users the service they need. Reliable systems are fault-tolerant: they expect faults and stop them from turning into failures. Some teams even inject faults on purpose, to prove the tolerance works before a real outage tests it.
| Kind of fault | Example | Main defence |
|---|---|---|
| Hardware | A disk dies, a machine loses power | Redundancy: replicas, spare machines, software that survives losing a node |
| Software | A bug that hits every node at once, a runaway process | Testing, isolating components, fast rollback, monitoring |
| Human | A bad config push, the wrong command in production | Safe places to experiment, gradual rollouts, easy undo, learning without blame |
Hardware faults tend to be random and independent. Software faults are the scarier kind, because the same bug runs on every machine at once. And people cause a large share of outages, which the chapter treats as a design problem rather than a discipline problem: make the right thing easy, the wrong thing hard, and mistakes cheap to undo. Blameless postmortems exist so that people report what happened honestly.
My take
Looking back, the script I wrote to copy production PostgreSQL data into dev and QA was a reliability tool. It gave people a realistic place to make mistakes that didn't matter, which is one of the chapter's main defences against human error.
Requirement 03
Scalability: grows how?
"Is it scalable?" is the wrong question. A better one is: if the load grows in this particular way, what are our options? That means first describing the load, for example requests per second, the ratio of reads to writes, the amount of data or the number of users online at once, and then asking which part breaks first.
| Architecture | What it means | Trade-off |
|---|---|---|
| Shared-memory (scale up) | A bigger machine | Simple, but cost climbs steeply and there's a ceiling |
| Shared-disk | Several machines, one shared storage system | Used by some warehouses, but contention limits how far it goes |
| Shared-nothing (scale out) | Independent machines, each with its own storage | Scales furthest, but now it's a distributed system (see Chapter 1) |
The advice is to split systems into parts that can grow independently, and not to build for scale you don't have yet. Architectures that fit one level of load rarely fit ten times that, so expect to rethink as you grow.
Requirement 04
Maintainability: the cost after launch
Most of the cost of software comes after it ships: fixing bugs, keeping it running, adapting it to new needs and paying down old decisions. The chapter splits maintainability into three goals:
- Operability: make it easy for the people running it to see what it's doing and keep it healthy.
- Simplicity: remove accidental complexity, the kind that comes from the implementation rather than the problem, mostly through good abstractions.
- Evolvability: make it easy to change when requirements change, which they will.
Push back
The timeline example is very web-app flavoured. From a data engineering seat the same ideas apply, since precomputing a timeline is really precomputing an aggregate, but you have to do that translation yourself.
Push back
The percentile advice is right, but measuring percentiles properly needs histogram tooling a small team may not have. A cheap place to start: log every request's or run's duration and compute the p95 once a day in SQL.
Push back
Maintainability gets the least concrete treatment, even though the chapter says it's where most of the cost is. Simplicity is hard to measure, and the section reads more like principles than practice.
Verdict
The most practical chapter so far
Chapter 1 gave me the vocabulary for architecture. Chapter 2 gives me the vocabulary for design reviews: p99, throughput vs response time, fault vs failure, operability. It's the chapter I'd hand to someone before their first on-call rotation.
Next up: Chapter 3, on data models and query languages.
Recap
Key takeaways
- Hello there! SAGUN made it to CHAPTER 2 of DDIA! This one is about how WELL a system works, not just what it does.
- Averages hide pain. Look at percentiles: the p99 is what your unluckiest users feel.
- One slow backend can make a whole request slow. The more services a request touches, the more the slow tail matters.
- A fault is one part breaking. A failure is the whole system letting users down. Good design stops faults from becoming failures.
- People cause a lot of outages. Give them safe places to test, quick rollbacks and good monitoring, not blame.
- Scalability isn't a yes or no. Ask: if load grows in THIS way, what are our options?
- Most of a system's cost comes after launch. Build it so others can run it, understand it and change it.
- SAGUN's DDIA level went up again! Next stop: CHAPTER 3!