← Reading DDIA

DDIA notes · Chapter 2

How well, not just what

Chapter 2 is about the requirements nobody writes in the ticket: how fast, how reliable, how big and how easy to change. It gives you the words to argue about them, and a few numbers to argue with.

Book
DDIA, 2nd edition
Authors
Martin Kleppmann & Chris Riccomini
Chapter
Defining Nonfunctional Requirements

The gist

The requirements nobody writes down

A feature request says what an app should do: show a timeline, send a message, run a report. It rarely says how fast, how often it may break, how much load it must take or how easy it must be to change later. Those are nonfunctional requirements, and Chapter 2 is about making them concrete enough to design for.

The chapter runs one example all the way through: a social network's home timeline, where opening the app shows recent posts from everyone you follow. You can build it by querying at read time, which gets expensive for people who follow thousands of accounts. Or you can precompute each user's timeline whenever someone posts, called fan-out on write, which makes reading cheap and posting expensive, especially for an account with millions of followers. Neither is free, and deciding between them is exactly what the rest of the chapter gives you tools for.

Requirement 01

Performance: look at the tail

Two numbers describe performance. Throughput is how much work gets done per second. Response time is how long one request takes from the user's point of view, including network delays and time spent waiting in a queue. The chapter keeps these apart from latency, the time a request spends waiting before anything works on it.

Response time isn't one number, because the same request can take 40 ms one moment and a second the next. So the chapter argues for percentiles instead of averages. The median (p50) is what a typical user sees. The p95, p99 and p999 are what your slowest requests look like, and those are often your heaviest users, the ones with the most data.

0500900fastestslowest40 requests, sorted · response time in msmean 128 msp50 79 msp95 380 ms
Fig. 1 — 40 illustrative response times, sorted. The mean is pulled up by a handful of slow requests, so it describes nobody; the percentiles say what the typical and the unlucky user actually get.

Why the tail gets worse

Two effects make slow requests matter more than their share suggests:

  • Tail latency amplification. If one page needs answers from ten backend calls, it's as slow as the slowest of the ten. A 1-in-100 slow call turns into a much more common slow page.
  • Queueing. A few slow requests can hold up the fast ones behind them, so measuring on the server alone can make things look better than users experience.

That's why service agreements are written in percentiles, like "p99 under 1 second", rather than averages.

My take

The async ETL pipeline I built was a throughput story, not a response-time one. It didn't make any single API call faster. It kept many calls in flight at once, so more wallets got processed per minute. Having separate words for the two made it easier to say what actually improved.

My take

My AWS Lambda monitoring reported whether each crawler run passed or failed. After this chapter I'd also track how long runs take, as percentiles. A run that succeeds but takes three times longer than usual is an early warning, and a pass/fail check can't see it.

Requirement 02

Reliability: faults aren't failures

The chapter draws a line I'll keep using: a fault is one part going wrong, like a disk, a process or a network link. A failure is the system as a whole no longer giving users the service they need. Reliable systems are fault-tolerant: they expect faults and stop them from turning into failures. Some teams even inject faults on purpose, to prove the tolerance works before a real outage tests it.

Kind of fault Example Main defence
Hardware A disk dies, a machine loses power Redundancy: replicas, spare machines, software that survives losing a node
Software A bug that hits every node at once, a runaway process Testing, isolating components, fast rollback, monitoring
Human A bad config push, the wrong command in production Safe places to experiment, gradual rollouts, easy undo, learning without blame

Hardware faults tend to be random and independent. Software faults are the scarier kind, because the same bug runs on every machine at once. And people cause a large share of outages, which the chapter treats as a design problem rather than a discipline problem: make the right thing easy, the wrong thing hard, and mistakes cheap to undo. Blameless postmortems exist so that people report what happened honestly.

My take

Looking back, the script I wrote to copy production PostgreSQL data into dev and QA was a reliability tool. It gave people a realistic place to make mistakes that didn't matter, which is one of the chapter's main defences against human error.

Requirement 03

Scalability: grows how?

"Is it scalable?" is the wrong question. A better one is: if the load grows in this particular way, what are our options? That means first describing the load, for example requests per second, the ratio of reads to writes, the amount of data or the number of users online at once, and then asking which part breaks first.

Architecture What it means Trade-off
Shared-memory (scale up) A bigger machine Simple, but cost climbs steeply and there's a ceiling
Shared-disk Several machines, one shared storage system Used by some warehouses, but contention limits how far it goes
Shared-nothing (scale out) Independent machines, each with its own storage Scales furthest, but now it's a distributed system (see Chapter 1)

The advice is to split systems into parts that can grow independently, and not to build for scale you don't have yet. Architectures that fit one level of load rarely fit ten times that, so expect to rethink as you grow.

Requirement 04

Maintainability: the cost after launch

Most of the cost of software comes after it ships: fixing bugs, keeping it running, adapting it to new needs and paying down old decisions. The chapter splits maintainability into three goals:

  • Operability: make it easy for the people running it to see what it's doing and keep it healthy.
  • Simplicity: remove accidental complexity, the kind that comes from the implementation rather than the problem, mostly through good abstractions.
  • Evolvability: make it easy to change when requirements change, which they will.

Push back

The timeline example is very web-app flavoured. From a data engineering seat the same ideas apply, since precomputing a timeline is really precomputing an aggregate, but you have to do that translation yourself.

Push back

The percentile advice is right, but measuring percentiles properly needs histogram tooling a small team may not have. A cheap place to start: log every request's or run's duration and compute the p95 once a day in SQL.

Push back

Maintainability gets the least concrete treatment, even though the chapter says it's where most of the cost is. Simplicity is hard to measure, and the section reads more like principles than practice.

Verdict

The most practical chapter so far

Chapter 1 gave me the vocabulary for architecture. Chapter 2 gives me the vocabulary for design reviews: p99, throughput vs response time, fault vs failure, operability. It's the chapter I'd hand to someone before their first on-call rotation.

Next up: Chapter 3, on data models and query languages.

Recap

Key takeaways

  1. Hello there! SAGUN made it to CHAPTER 2 of DDIA! This one is about how WELL a system works, not just what it does.
  2. Averages hide pain. Look at percentiles: the p99 is what your unluckiest users feel.
  3. One slow backend can make a whole request slow. The more services a request touches, the more the slow tail matters.
  4. A fault is one part breaking. A failure is the whole system letting users down. Good design stops faults from becoming failures.
  5. People cause a lot of outages. Give them safe places to test, quick rollbacks and good monitoring, not blame.
  6. Scalability isn't a yes or no. Ask: if load grows in THIS way, what are our options?
  7. Most of a system's cost comes after launch. Build it so others can run it, understand it and change it.
  8. SAGUN's DDIA level went up again! Next stop: CHAPTER 3!