The System Design Mindset
Master core system design foundations: trade-offs, data operations, availability, SLOs/SLAs, reliability, throughput, latency, and scaling strategies.
Lesson goal
By the end of this lesson, you will understand how to evaluate architectural trade-offs, trace data lifecycle operations, measure availability and service levels (SLOs/SLAs), and address reliability, throughput, and latency through scaling strategies.
System design is a process of making thoughtful choices about how a system should behave. It is not about finding a perfect solution; it is about analyzing trade-offs and selecting the best available option.
The system design mindset
Every architectural decision comes with advantages and compromises. Effective system design follows a practical three-step process:
- Analyze trade-offs — Understand what each possible choice improves and what it may compromise.
- Improve system behavior — Use those insights to make the system work better.
- Choose the best available option — Select the option that fits the situation most effectively, even when every option has limitations.
Key idea
Good system design accepts that no solution is perfect. The goal is to make informed decisions that improve the system overall.
Moving, storing, and transforming data
Large applications ultimately deal with data in three fundamental ways: they move it, store it, and transform it. These operations take place across machines, databases, blob stores, file systems, and distributed systems.
The three operations
| Operation | Focus |
|---|---|
| Moving data | Data travels across the machines and systems that make up an application. |
| Storing data | Data is kept in databases, blob stores, or file systems. |
| Transforming data | Data is processed as it moves through an application or distributed system. |
Availability, SLOs, and SLAs
Availability
Availability describes how often a service is operational and usable. It is commonly expressed as a percentage:
Availability = (Time the service is available / Total measured time) × 100A higher percentage means the service is available for more of the measured period.
Percentages and the "nines" convention
Availability percentages are often described by the number of 9s they contain:
| Percentage | Common description | Allowed downtime per year |
|---|---|---|
| 99% | Two nines | ~3.65 days / year |
| 99.9% | Three nines | ~8.76 hours / year |
| 99.99% | Four nines | ~52.56 minutes / year |
| 99.999% | Five nines | ~5.26 minutes / year |

Github availability graph for 19.08.2026 - source 'https://githubstatus.com'
The nines convention gives a compact way to describe an availability target. Even a small increase in the percentage represents a substantially stricter target requiring greater architectural resilience.
SLOs and SLAs
A service level objective (SLO) is a target for a service level, such as its availability percentage. It states the level the service aims to provide.
A service level agreement (SLA) is a formal agreement that defines the expected service level, making the expectation explicit between the service provider and its customers (often including financial penalties for downtime).
The distinction is straightforward:
- SLO: The internal or external target for the service level.
- SLA: The formal agreement that defines the expected service level.
Recap: Availability & Service Levels
Availability measures how often a service is usable. Percentages are expressed using the "nines" convention. An SLO is a service-level target, while an SLA is a formal agreement defining the expected service level.

Reliability, throughput, and latency
Designing a system means deciding how well it must continue working, how much work it must handle, and how quickly it must respond. Three central requirements are reliability, throughput, and latency.
Core concepts
- Reliability — The ability of a system to operate correctly as expected.
- Fault tolerance — The ability of a system to continue operating when faults occur.
- Redundancy — Having additional duplicate components or resources that can take over when needed.
- Throughput — The amount of work or number of requests a system processes over a period of time (e.g., requests per second).
- Latency — The duration required for a system to respond to a request or complete an operation.
Reliability is closely related to fault tolerance and redundancy: a system becomes more reliable when it can tolerate component faults and has redundant resources available.
Throughput and latency describe different aspects of performance:
- A system can process a massive volume of work (high throughput) while individual operations still take noticeable time (higher latency).
- Conversely, a single operation can execute instantly (low latency) without the system being capable of handling heavy concurrent volume (low throughput).
Scaling to meet requirements
Scaling changes a system's capacity so it can better satisfy reliability, throughput, or latency requirements.
| Approach | Description | Requirements addressed |
|---|---|---|
| Vertical scaling | Increase compute resources (CPU, RAM, disk) of a single server. | Throughput or latency handled within a single node's capacity limits |
| Horizontal scaling | Add more servers and distribute workload across them. | Throughput, reliability, and fault tolerance across nodes |
Vertical scaling
Vertical scaling focuses on a single machine. Increasing its resources helps it process more work or complete operations faster, but remains bounded by hardware limits and represents a single point of failure.
Horizontal scaling
Horizontal scaling adds more machines instead of relying on one. Additional nodes increase total throughput capacity and provide redundancy, supporting fault tolerance and high availability.

Keep the requirements distinct
Reliability, throughput, and latency are separate requirements. Fault tolerance and redundancy support reliability; scaling increases throughput.
Key takeaways
- Thoughtful trade-offs: System design is about analyzing trade-offs and choosing the best available option for your specific constraints.
- Data lifecycle: Large applications continuously move, store, and transform data across machines, databases, and compute workers.
- Availability: Availability measures uptime percentage.
- Performance dimensions: Reliability ensures correctness and uptime, throughput measures volume of work, and latency measures response time.
- Scaling strategies: Use vertical scaling for simple compute boosts, horizontal scaling for scalable throughput and redundancy.