The gist
Four questions, no defaults
Most applications today are limited less by CPU than by data: how much there is, how complex it is, and how fast it changes. Those applications are built from the same few building blocks: databases, caches, search indexes, stream processors and batch jobs. Chapter 1 doesn't explain how any of them work. It explains the decisions that come before choosing them, and it frames each one as a spectrum rather than a right answer.
Trade-off 01
Operational vs analytical
The first split is between systems that run the business and systems that study it. An operational database handles a checkout or a login: small reads and writes, one record at a time, and it needs to be fast now. An analytical system answers "how did sales change last quarter?" by scanning millions of rows, often from a copy of the data.
| Operational (OLTP) | Analytical (OLAP) | |
|---|---|---|
| Reads | Look up a few records by key | Aggregate over huge numbers of records |
| Writes | Create, update, delete single records | Bulk loads or a stream of events |
| Used by | Customers, through the app | Analysts and data scientists |
| Queries | A fixed set, written by developers | Anything someone thinks to ask |
| Data shows | The current state | The history of what happened |
Because the two workloads fight each other, companies copy data out of operational databases into a data warehouse through ETL (extract, transform, load), or dump it raw into a data lake and decide later what it means. The chapter then introduces the idea I expect to see in every later chapter: a system of record holds the authoritative version of the data, and everything else, such as caches, indexes, warehouses and models, is derived data that can be rebuilt from it.
My take
"Which system is the source of truth?" is the most useful question in the chapter,
and I wish I'd asked it earlier. In ScamFilter, the
system of record was really the blockchain, read through Etherscan. My PostgreSQL
transactions table, the spam_detections table and the wallet metrics in
MongoDB were all derived. Seeing it that way changes how you treat a bug: if derived
data is wrong, you don't patch rows by hand, you fix the logic and rebuild.
Push back
The OLTP/OLAP table is tidy, but real systems blur it, and the chapter admits as much. Product features like recommendations and in-app dashboards run analytical queries inside the operational app, and "reverse ETL" pushes warehouse results back into production tools. Treat the table as a starting point, not a rule.
Trade-off 02
Cloud vs self-hosting
This is the classic build-or-buy decision. A managed cloud service means you don't run the servers, capacity can grow and shrink with demand, and specialists handle the hard parts. The price is control: you can't fix a bug in the service yourself, you see less of what's happening inside it when something goes wrong, you depend on the vendor staying around and staying affordable, and a steady, predictable workload can cost more rented than owned.
The chapter also describes what "cloud-native" changes about database design. The big shift is separating storage from compute: data lives in cheap, durable object storage like S3, and the machines that run queries can be added or removed independently. And operations work doesn't disappear. It moves from patching machines to choosing services, wiring them together and watching the bill.
My take
That last point matches my experience. I built pipeline monitoring on AWS Lambda, so there were no servers to look after, but I still had to decide what counted as a failure, log status reports and send alerts to Discord when a crawler broke. Serverless removed the machines. It didn't remove the need to know when things go wrong.
Push back
The chapter's honest answer is "it depends", which can feel unsatisfying. What makes it useful anyway is the list of what it depends on: team size, how predictable the load is, how much control you need and how much lock-in you can live with. I'd use that list as a checklist before choosing.
Trade-off 03
Single node vs distributed
There are good reasons to spread a system across machines: the users are spread out, one machine can't hold the data or handle the load, you need to survive a machine failing, or the law says certain data must stay in a certain country. The chapter also covers microservices and serverless as distributed designs.
The cost is steep. Every call over the network can be slow or fail in ways a function call on the same machine can't. Finding a bug means following one request across many services, which is why tracing and observability tools exist. And keeping data consistent across services becomes your problem. The chapter's advice is plain: if a single machine can do the job, it's usually simpler and cheaper.
My take
This is the advice I'd most like to hear repeated. Single machines are powerful now: ScamFilter's metrics engine runs in Polars on one machine, and for its workload that was the right call. Distribution is something you should need before you reach for it, not a sign of a serious system.
Trade-off 04
Data, law and society
The chapter ends somewhere most database books never go: the law. Rules like the EU's GDPR give people the right to have their data deleted, which is awkward when your system is built on append-only logs and immutable files. The chapter argues for data minimisation: collect and keep only what you need, because stored data is a liability as well as an asset.
Push back
The chapter holds two ideas in tension without fully resolving them. Earlier it describes the data-lake habit of keeping raw data because you might need it later. Here it says to keep only what you need. Both are reasonable, and in practice the fix is to decide retention up front, per dataset, instead of defaulting to "keep everything".
Blockchain data makes this interesting. On-chain transactions are public and can't be deleted. A wallet profile built from them can still point to a real person, and the analytics about the wallet can be deleted. The raw data may be public, but what you derive from it is still yours to answer for.
Verdict
Worth reading, but not for internals
Chapter 1 is a map, not a manual. If you want to know how a storage engine works, that comes later, and this chapter can feel like a long list of terms, each introduced briefly. What it does well is give you the vocabulary for architecture discussions, and a habit worth keeping: when someone says a design is "best", ask best at what, and at what cost.
Next up: Chapter 2, on defining nonfunctional requirements such as reliability, scalability and maintainability.
Recap
Key takeaways
- Hello there! Welcome to the world of DATA SYSTEMS! My name is OAK. SAGUN just finished chapter 1 of DDIA. Here's what stuck.
- There is no best database. There's only best at what, and at what cost.
- First, find the system of record. Caches, indexes and warehouses are derived data, and derived data can be rebuilt.
- Operational and analytical systems want different things, so they usually live apart. Real apps blur the line, though!
- The cloud removes the servers, not the operations. Someone still has to know when things break.
- If one machine can do the job, use one machine. Every network call is a new way to fail.
- Data is a liability as well as an asset. Decide how long to keep it before you collect it.
- SAGUN's DDIA level went up! Next stop: CHAPTER 2!