Engineered
Application Architecture

The System Design Mindset

Master core system design foundations: trade-offs, data operations, availability, SLOs/SLAs, reliability, throughput, latency, and scaling strategies.

Lesson goal

By the end of this lesson, you will understand how to evaluate architectural trade-offs, trace data lifecycle operations, measure availability and service levels (SLOs/SLAs), and address reliability, throughput, and latency through scaling strategies.

System design is a process of making thoughtful choices about how a system should behave. It is not about finding a perfect solution; it is about analyzing trade-offs and selecting the best available option.

The system design mindset

Every architectural decision comes with advantages and compromises. Effective system design follows a practical three-step process:

  1. Analyze trade-offs — Understand what each possible choice improves and what it may compromise.
  2. Improve system behavior — Use those insights to make the system work better.
  3. Choose the best available option — Select the option that fits the situation most effectively, even when every option has limitations.

Key idea

Good system design accepts that no solution is perfect. The goal is to make informed decisions that improve the system overall.

Moving, storing, and transforming data

Large applications ultimately deal with data in three fundamental ways: they move it, store it, and transform it. These operations take place across machines, databases, blob stores, file systems, and distributed systems.

The three operations

OperationFocus
Moving dataData travels across the machines and systems that make up an application.
Storing dataData is kept in databases, blob stores, or file systems.
Transforming dataData is processed as it moves through an application or distributed system.

Availability, SLOs, and SLAs

Availability

Availability describes how often a service is operational and usable. It is commonly expressed as a percentage:

Availability = (Time the service is available / Total measured time) × 100

A higher percentage means the service is available for more of the measured period.

Percentages and the "nines" convention

Availability percentages are often described by the number of 9s they contain:

PercentageCommon descriptionAllowed downtime per year
99%Two nines~3.65 days / year
99.9%Three nines~8.76 hours / year
99.99%Four nines~52.56 minutes / year
99.999%Five nines~5.26 minutes / year

Github availability graph for 19.08.2026 - source 'https://githubstatus.com'

The nines convention gives a compact way to describe an availability target. Even a small increase in the percentage represents a substantially stricter target requiring greater architectural resilience.

SLOs and SLAs

A service level objective (SLO) is a target for a service level, such as its availability percentage. It states the level the service aims to provide.

A service level agreement (SLA) is a formal agreement that defines the expected service level, making the expectation explicit between the service provider and its customers (often including financial penalties for downtime).

The distinction is straightforward:

  • SLO: The internal or external target for the service level.
  • SLA: The formal agreement that defines the expected service level.

Recap: Availability & Service Levels

Availability measures how often a service is usable. Percentages are expressed using the "nines" convention. An SLO is a service-level target, while an SLA is a formal agreement defining the expected service level.


Reliability, throughput, and latency

Designing a system means deciding how well it must continue working, how much work it must handle, and how quickly it must respond. Three central requirements are reliability, throughput, and latency.

Core concepts

  • Reliability — The ability of a system to operate correctly as expected.
  • Fault tolerance — The ability of a system to continue operating when faults occur.
  • Redundancy — Having additional duplicate components or resources that can take over when needed.
  • Throughput — The amount of work or number of requests a system processes over a period of time (e.g., requests per second).
  • Latency — The duration required for a system to respond to a request or complete an operation.

Reliability is closely related to fault tolerance and redundancy: a system becomes more reliable when it can tolerate component faults and has redundant resources available.

Throughput and latency describe different aspects of performance:

  • A system can process a massive volume of work (high throughput) while individual operations still take noticeable time (higher latency).
  • Conversely, a single operation can execute instantly (low latency) without the system being capable of handling heavy concurrent volume (low throughput).

Scaling to meet requirements

Scaling changes a system's capacity so it can better satisfy reliability, throughput, or latency requirements.

ApproachDescriptionRequirements addressed
Vertical scalingIncrease compute resources (CPU, RAM, disk) of a single server.Throughput or latency handled within a single node's capacity limits
Horizontal scalingAdd more servers and distribute workload across them.Throughput, reliability, and fault tolerance across nodes

Vertical scaling

Vertical scaling focuses on a single machine. Increasing its resources helps it process more work or complete operations faster, but remains bounded by hardware limits and represents a single point of failure.

Horizontal scaling

Horizontal scaling adds more machines instead of relying on one. Additional nodes increase total throughput capacity and provide redundancy, supporting fault tolerance and high availability.

Keep the requirements distinct

Reliability, throughput, and latency are separate requirements. Fault tolerance and redundancy support reliability; scaling increases throughput.

Key takeaways

  • Thoughtful trade-offs: System design is about analyzing trade-offs and choosing the best available option for your specific constraints.
  • Data lifecycle: Large applications continuously move, store, and transform data across machines, databases, and compute workers.
  • Availability: Availability measures uptime percentage.
  • Performance dimensions: Reliability ensures correctness and uptime, throughput measures volume of work, and latency measures response time.
  • Scaling strategies: Use vertical scaling for simple compute boosts, horizontal scaling for scalable throughput and redundancy.

How is this lesson?