← Reading DDIA

DDIA notes · Chapter 1

Every system is a trade-off

The first chapter of Designing Data-Intensive Applications barely mentions a database. There's no B-tree and no replication log. Instead it asks four questions you have to answer before you pick any tool, and it refuses to give you default answers.

Book
DDIA, 2nd edition
Authors
Martin Kleppmann & Chris Riccomini
Chapter
Trade-offs in Data Systems Architecture

The gist

Four questions, no defaults

Most applications today are limited less by CPU than by data: how much there is, how complex it is, and how fast it changes. Those applications are built from the same few building blocks: databases, caches, search indexes, stream processors and batch jobs. Chapter 1 doesn't explain how any of them work. It explains the decisions that come before choosing them, and it frames each one as a spectrum rather than a right answer.

OPERATIONAL ANALYTICAL SELF-HOSTED CLOUD SINGLE NODE DISTRIBUTED KEEP EVERYTHING KEEP ONLY WHAT YOU NEED ? ? ? ?
Fig. 1 — The chapter's four trade-offs. Where a system sits on each line depends on what it's for.

Trade-off 01

Operational vs analytical

The first split is between systems that run the business and systems that study it. An operational database handles a checkout or a login: small reads and writes, one record at a time, and it needs to be fast now. An analytical system answers "how did sales change last quarter?" by scanning millions of rows, often from a copy of the data.

Operational (OLTP)Analytical (OLAP)
ReadsLook up a few records by keyAggregate over huge numbers of records
WritesCreate, update, delete single recordsBulk loads or a stream of events
Used byCustomers, through the appAnalysts and data scientists
QueriesA fixed set, written by developersAnything someone thinks to ask
Data showsThe current stateThe history of what happened

Because the two workloads fight each other, companies copy data out of operational databases into a data warehouse through ETL (extract, transform, load), or dump it raw into a data lake and decide later what it means. The chapter then introduces the idea I expect to see in every later chapter: a system of record holds the authoritative version of the data, and everything else, such as caches, indexes, warehouses and models, is derived data that can be rebuilt from it.

My take

"Which system is the source of truth?" is the most useful question in the chapter, and I wish I'd asked it earlier. In ScamFilter, the system of record was really the blockchain, read through Etherscan. My PostgreSQL transactions table, the spam_detections table and the wallet metrics in MongoDB were all derived. Seeing it that way changes how you treat a bug: if derived data is wrong, you don't patch rows by hand, you fix the logic and rebuild.

Push back

The OLTP/OLAP table is tidy, but real systems blur it, and the chapter admits as much. Product features like recommendations and in-app dashboards run analytical queries inside the operational app, and "reverse ETL" pushes warehouse results back into production tools. Treat the table as a starting point, not a rule.

Trade-off 02

Cloud vs self-hosting

This is the classic build-or-buy decision. A managed cloud service means you don't run the servers, capacity can grow and shrink with demand, and specialists handle the hard parts. The price is control: you can't fix a bug in the service yourself, you see less of what's happening inside it when something goes wrong, you depend on the vendor staying around and staying affordable, and a steady, predictable workload can cost more rented than owned.

The chapter also describes what "cloud-native" changes about database design. The big shift is separating storage from compute: data lives in cheap, durable object storage like S3, and the machines that run queries can be added or removed independently. And operations work doesn't disappear. It moves from patching machines to choosing services, wiring them together and watching the bill.

My take

That last point matches my experience. I built pipeline monitoring on AWS Lambda, so there were no servers to look after, but I still had to decide what counted as a failure, log status reports and send alerts to Discord when a crawler broke. Serverless removed the machines. It didn't remove the need to know when things go wrong.

Push back

The chapter's honest answer is "it depends", which can feel unsatisfying. What makes it useful anyway is the list of what it depends on: team size, how predictable the load is, how much control you need and how much lock-in you can live with. I'd use that list as a checklist before choosing.

Trade-off 03

Single node vs distributed

There are good reasons to spread a system across machines: the users are spread out, one machine can't hold the data or handle the load, you need to survive a machine failing, or the law says certain data must stay in a certain country. The chapter also covers microservices and serverless as distributed designs.

The cost is steep. Every call over the network can be slow or fail in ways a function call on the same machine can't. Finding a bug means following one request across many services, which is why tracing and observability tools exist. And keeping data consistent across services becomes your problem. The chapter's advice is plain: if a single machine can do the job, it's usually simpler and cheaper.

My take

This is the advice I'd most like to hear repeated. Single machines are powerful now: ScamFilter's metrics engine runs in Polars on one machine, and for its workload that was the right call. Distribution is something you should need before you reach for it, not a sign of a serious system.

Trade-off 04

Data, law and society

The chapter ends somewhere most database books never go: the law. Rules like the EU's GDPR give people the right to have their data deleted, which is awkward when your system is built on append-only logs and immutable files. The chapter argues for data minimisation: collect and keep only what you need, because stored data is a liability as well as an asset.

Push back

The chapter holds two ideas in tension without fully resolving them. Earlier it describes the data-lake habit of keeping raw data because you might need it later. Here it says to keep only what you need. Both are reasonable, and in practice the fix is to decide retention up front, per dataset, instead of defaulting to "keep everything".

Blockchain data makes this interesting. On-chain transactions are public and can't be deleted. A wallet profile built from them can still point to a real person, and the analytics about the wallet can be deleted. The raw data may be public, but what you derive from it is still yours to answer for.

Verdict

Worth reading, but not for internals

Chapter 1 is a map, not a manual. If you want to know how a storage engine works, that comes later, and this chapter can feel like a long list of terms, each introduced briefly. What it does well is give you the vocabulary for architecture discussions, and a habit worth keeping: when someone says a design is "best", ask best at what, and at what cost.

Next up: Chapter 2, on defining nonfunctional requirements such as reliability, scalability and maintainability.

Recap

Key takeaways

  1. Hello there! Welcome to the world of DATA SYSTEMS! My name is OAK. SAGUN just finished chapter 1 of DDIA. Here's what stuck.
  2. There is no best database. There's only best at what, and at what cost.
  3. First, find the system of record. Caches, indexes and warehouses are derived data, and derived data can be rebuilt.
  4. Operational and analytical systems want different things, so they usually live apart. Real apps blur the line, though!
  5. The cloud removes the servers, not the operations. Someone still has to know when things break.
  6. If one machine can do the job, use one machine. Every network call is a new way to fail.
  7. Data is a liability as well as an asset. Decide how long to keep it before you collect it.
  8. SAGUN's DDIA level went up! Next stop: CHAPTER 2!