ersantana.com, mapped
A JSON Canvas rendered by Quartz 5. File nodes link to real notes; hover them for a preview.
Partition Tolerance in CAP Theorem: The Inevitable Necessity
Understanding Partition Tolerance (P)
Partition Tolerance (P) in the CAP theorem means that a distributed system continues to operate even if there are arbitrary message losses or failures within the network that separate the system into multiple isolated partitions. Imagine your distributed database servers are connected by a network. A “network partition” means that communication between some of these servers is lost, effectively splitting them into two or more independent groups that cannot talk to each other. (Gilbert & Lynch, 2002)
A partition-tolerant system is designed to handle these communication breakdowns gracefully. It doesn’t simply crash or become unusable. Instead, it makes a choice: it can either sacrifice consistency to remain available within each partition, or it can sacrifice availability to guarantee consistency across all partitions. The key takeaway is that network partitions are an inevitable reality in any sufficiently large or complex distributed system.
Why Partition Tolerance is a Must-Have
In the real world, networks are unreliable. Cables get cut, routers fail, switches misbehave, cloud availability zones experience outages, and even software bugs can lead to communication issues. These events are not rare anomalies; they are expected occurrences in any system that spans multiple machines, data centers, or geographical regions.
Consider a system without partition tolerance. If a network partition occurs, and your system isn’t designed to handle it, the default behavior would likely be a complete system halt or significant data corruption. If two parts of your system can no longer communicate, they can’t agree on the state of the data. If both parts continue to operate independently without a strategy for partition tolerance, they will diverge, leading to inconsistent data or outright data loss when the partition heals.
Therefore, any robust distributed system must be partition tolerant to function reliably in a real-world network environment. It’s not a choice but a fundamental requirement for survivability. This is the point Brewer himself made when revisiting the theorem twelve years on, and the one Kleppmann’s critique sharpens: the only real choice is how the system behaves during a partition (Brewer, 2012; Kleppmann, 2015). Abadi’s PACELC formulation adds the latency trade-off that applies even when the network is healthy (Abadi, 2012).
Real-World Examples of Partition Tolerance in Action
Because Partition Tolerance is a given, these examples highlight the trade-off made between Consistency (C) and Availability (A) when a partition occurs:
Eventually Consistent NoSQL Databases (AP Choice)
Example: Cassandra, DynamoDB
These databases are designed for high availability and partition tolerance (AP). During a network partition, nodes in different partitions continue to accept writes and serve reads. This means that if a write happens in one partition, and a read happens in another isolated partition, the read might receive stale data. However, the system remains available. When the partition heals, the data is asynchronously synchronized across the nodes, eventually reaching a consistent state.
Use Case: Social media feeds, IoT data collection, content delivery networks Trade-off: Temporary inconsistency for continuous operation
Distributed Relational Databases with Strong Consistency (CP Choice)
Example: CockroachDB, Google Cloud Spanner
These systems often prioritize consistency and partition tolerance (CP). When a network partition occurs, if a node cannot communicate with the primary or quorum of other nodes needed to guarantee consistency, it will typically become unavailable for writes (and sometimes reads) in that partition. This ensures that no inconsistent data is written.
Use Case: Financial systems, banking applications, legal record systems Trade-off: Temporary unavailability for data integrity
Google Drive/Dropbox Offline Editing (AP Choice)
When you’re editing a document in Google Drive offline, you’re essentially in a “partition” from the main cloud service. The system allows you to continue working (maintaining availability), but your changes aren’t immediately synchronized with the cloud or other collaborators (sacrificing immediate consistency). When your internet connection is restored (the partition heals), the changes are synced, and conflict resolution mechanisms come into play to eventually achieve consistency.
Trade-off: Potential conflicts for continued productivity
Distributed Caching Systems (Configurable)
Example: Redis Cluster
Some caching systems can operate in AP mode. If a partition occurs, nodes might continue to serve cached data even if it’s not the absolute latest from the source of truth, to maintain availability. Other configurations might prioritize CP, making parts of the cache unavailable if they cannot guarantee consistent access to the underlying data source.
Trade-off: Depends on configuration - either stale cache data or temporary cache unavailability
Microservices Architectures
In a well-designed microservices architecture, individual services can be deployed and scaled independently. If one service experiences a network partition (e.g., a database service becomes isolated), other services that don’t depend on it can continue to function (availability). Services that do depend on the partitioned service must then decide whether to:
- Fail requests (prioritizing consistency by not providing potentially incorrect data)
- Return fallback or cached responses (prioritizing availability)
This highlights the architectural pattern’s built-in partition tolerance, pushing the C/A trade-off to individual service design.
Content Delivery Networks (CDNs) - AP Choice
CDNs are distributed by nature, with edge servers around the globe. When a regional partition occurs:
- Available Response: Edge servers continue serving cached content even if they can’t communicate with origin servers
- Eventual Consistency: New content updates propagate when connectivity is restored
- User Experience: Websites remain accessible, even with potentially slightly stale content
Database Choices for Partition Tolerance
Since Partition Tolerance is mandatory for distributed systems, the choice focuses on how the database handles partitions by prioritizing either Consistency (C) or Availability (A).
For CP Systems: CockroachDB
Why CockroachDB excels for CP partition handling:
- Distributed Consensus: Uses Raft protocol to ensure all committed transactions are truly consistent across the cluster
- Quorum Requirements: If a node cannot form a quorum with its peers during a partition, it stops processing writes/reads until consistency can be re-established
- Geographic Distribution: Provides strong consistency even across globally distributed data
- Automatic Recovery: When partitions heal, the system automatically resumes normal operations
Ideal for:
- Multi-region banking systems
- Enterprise resource planning (ERP) systems
- Any application where data accuracy is non-negotiable
Partition Behavior: Becomes unavailable in affected partitions to guarantee no inconsistent data
For AP Systems: Apache Cassandra
Why Cassandra excels for AP partition handling:
- Masterless Architecture: Every node can accept writes and serve reads independently
- Continued Operation: Accepts writes and serves reads within each isolated partition during network splits
- Eventual Consistency: Has mechanisms (anti-entropy via Merkle trees, read repairs) to reconcile divergent data when partitions heal
- Tunable Consistency: Can adjust consistency levels based on requirements
Ideal for:
- Large-scale IoT data pipelines
- Social media message queues
- Gaming leaderboards
- Any system requiring continuous operation
Partition Behavior: Continues operating in each partition, reconciling data when connectivity restores
Hybrid Approach: Different Services, Different Choices
Many real-world systems use different approaches for different components:
E-commerce Platform Example:
- Product Catalog: Cassandra (AP) - browsing must always work
- User Authentication: PostgreSQL (CP) - account security is critical
- Shopping Cart: Redis (AP) - session continuity is important
- Payment Processing: Traditional RDBMS (CP) - financial accuracy is paramount
- Recommendation Engine: Cassandra (AP) - suggestions can be eventually consistent
Key Principles for Partition Tolerance
- Expect Partitions: Design systems assuming network failures will occur
- Make Explicit Choices: Decide between consistency and availability for each component
- Plan for Recovery: Have mechanisms to reconcile data when partitions heal
- Monitor and Alert: Implement monitoring to detect and respond to partitions
- Test Regularly: Use chaos engineering to test partition scenarios
Conclusion
Partition tolerance is not an optional feature but a foundational requirement for any realistic distributed system. The CAP theorem simply states that when these inevitable partitions occur, you must choose between maintaining full consistency or ensuring continuous availability for your clients.
The key insight is that this choice should be made deliberately based on business requirements rather than by accident. Understanding how your chosen database and architecture handle partitions is crucial for building reliable distributed systems that behave predictably under adverse network conditions.
CP Systems: Prioritizing Consistency Over Availability During Partitions
Understanding CP Systems
When a distributed system must handle network partitions (which is an unavoidable reality in any non-trivial distributed system), a CP system chooses to prioritize Consistency (C) over Availability (A) during such an event. This means that if a network partition occurs, and a part of the system cannot communicate with the rest to guarantee that data is consistent across all nodes, that part of the system will become unavailable.
The system will either block operations, refuse to serve requests, or return an error rather than risk providing stale or incorrect data. The core principle here is that data integrity and accuracy are paramount. No matter what, if you read data, you are guaranteed to get the latest, most accurate version, or you’ll get an error telling you the system cannot currently fulfill that guarantee.
CP systems often use consensus algorithms (like Paxos or Raft) to ensure that all committed writes are agreed upon by a majority of nodes before being acknowledged. Raft in particular was designed so that this agreement protocol could be understood and implemented correctly by ordinary engineering teams (Ongaro & Ousterhout, 2014).
Key Characteristics of CP Systems
- Strong Consistency Guarantees: All reads return the most recent write or an error
- Consensus-Based Operations: Use algorithms like Raft or Paxos for agreement
- Quorum Requirements: Operations require majority node agreement
- Partition Response: Become unavailable rather than serve inconsistent data
- ACID Compliance: Often support full database transactions
- Synchronous Replication: Changes must be confirmed across replicas before acknowledgment
What a quorum write looks like
sequenceDiagram participant C as Client participant L as Leader participant F1 as Follower 1 participant F2 as Follower 2 C->>L: write(x = 42) L->>F1: append entry L->>F2: append entry F1-->>L: ack Note over L: majority reached (2 of 3) L-->>C: committed F2-->>L: ack (late, still fine)
The price of consistency
If the leader cannot reach a majority, it must refuse the write. During a partition that isolates the leader, a CP system returns errors or times out rather than accepting data it cannot replicate. Design the client for that: retries with backoff, idempotent writes, and a clear user-facing failure state.
Real-World Examples of CP Systems in Action
Financial Transaction Systems
Scenario: A bank’s core ledger that records account balances and transactions.
It’s absolutely critical that every read of an account balance reflects the absolute latest state. If a network partition occurs and a subset of the bank’s servers can’t confirm the latest transactions with the main cluster, those servers will halt operations or become read-only, refusing to process new debits or credits.
Why CP is Essential:
- Prevents double-spending or lost transactions
- Ensures regulatory compliance
- Maintains customer trust
- Avoids financial discrepancies
Trade-off: Temporary unavailability in partitioned regions vs. potential financial errors
Distributed Locking Services
Examples: Apache ZooKeeper, etcd
These systems are used by other distributed applications to manage shared configurations, name services, and crucial distributed locks. When an application needs to acquire a lock to perform a critical operation (like modifying a unique resource), it’s essential that only one application holds that lock globally.
Partition Behavior: If a network partition prevents a ZooKeeper or etcd node from reaching a quorum of its peers to confirm the lock’s state, it will refuse to grant new locks or even become unavailable for reads until the partition heals.
Why CP is Essential:
- Prevents multiple applications from simultaneously thinking they have the same lock
- Avoids severe data corruption from concurrent modifications
- Maintains system-wide coordination
Distributed SQL Databases
Examples: CockroachDB, Google Cloud Spanner
These “NewSQL” databases aim to provide the strong consistency guarantees of traditional SQL databases in a distributed, scalable environment. When a network partition occurs, they ensure that transactions maintain full ACID properties.
Partition Behavior: If a node cannot communicate with a sufficient number of its replicas to establish a “quorum” for a write operation, it will pause operations or become temporarily unavailable for writes until the quorum can be re-established.
Why CP is Essential:
- Maintains ACID transaction guarantees
- Ensures referential integrity across tables
- Supports complex business logic requiring consistency
Customer Order Processing Systems
Scenario: Multi-step order processing system
When an order moves from “payment received” to “items picked,” it’s vital that all parts of the system (inventory, shipping, customer service) consistently see the correct, single state of that order.
Partition Behavior: If a partition prevents synchronization, the system might block the order from progressing, ensuring that an item isn’t shipped twice or an incorrect payment status is displayed.
Trade-off: Momentary processing delays vs. incorrect order fulfillment
Healthcare Record Systems
Scenario: Critical patient data management
For critical patient records (e.g., medication orders, life-support settings), absolute consistency is paramount. If a network partition means a physician’s workstation cannot retrieve the confirmed, latest version of a patient’s medication list from all synchronized replicas, the system should prevent the physician from proceeding with a new order.
Why CP is Critical:
- Prevents medical errors from inconsistent data
- Ensures patient safety
- Maintains treatment continuity
- Supports regulatory compliance
Trade-off: Temporary system unavailability vs. potential medical mistakes
Database Choices for CP Systems
Primary Recommendation: CockroachDB
Why CockroachDB is ideal for CP systems:
- Distributed SQL: Provides familiar SQL interface with distributed consistency
- Raft Consensus: Uses Raft protocol for strong consistency across nodes
- ACID Transactions: Full support for complex, multi-table transactions
- Automatic Partitioning: Handles data distribution while maintaining consistency
- Global Consistency: Ensures consistency even across geographic regions
- Partition Handling: Stops operations in affected partitions until consistency can be guaranteed
Key Features:
- Serializable Isolation: Strongest consistency level available
- Multi-Version Concurrency Control (MVCC): Handles concurrent operations safely
- Automatic Rebalancing: Maintains optimal data distribution
- Built-in Fault Detection: Quickly identifies and responds to failures
Ideal Use Cases:
- Multi-region banking systems
- Enterprise resource planning (ERP) systems
- E-commerce transaction processing
- Any application requiring global data consistency
Alternative Options
Google Cloud Spanner:
- Globally distributed SQL database
- TrueTime API for global consistency
- Automatic scaling with consistent performance
- Ideal for large-scale enterprise applications
Traditional RDBMS with Synchronous Replication:
- PostgreSQL: With synchronous replication and strong isolation
- MySQL: With synchronous replication configured
- Oracle RAC: For enterprise-scale consistent operations
Apache Cassandra (when configured for CP):
- Can be configured with strong consistency levels
- Quorum reads and writes
- Less common configuration but possible for specific use cases
etcd:
- Distributed key-value store
- Built on Raft consensus
- Primarily for configuration management and service discovery
Implementation Considerations
Consensus Algorithm Choice
Raft Protocol:
- Easier to understand and implement
- Clear leader election process
- Used by etcd, CockroachDB
Paxos Protocol:
- More complex but highly proven
- Better for some specific scenarios
- Used by Google’s systems
Quorum Configuration
Simple Majority:
- Requires (N/2 + 1) nodes for operations
- Good balance of consistency and availability
All Nodes:
- Requires all nodes to agree
- Highest consistency, lowest availability
Configurable Quorums:
- Different requirements for reads vs. writes
- Tunable based on specific needs
Monitoring and Alerting
Key Metrics to Monitor:
- Partition detection and duration
- Quorum status and health
- Transaction latency and throughput
- Node availability and connectivity
- Consistency lag metrics
Trade-offs and Considerations
Benefits of CP Systems
- Data Integrity: Guaranteed consistent data across all nodes
- Regulatory Compliance: Meets strict consistency requirements
- Simplified Application Logic: No need to handle eventual consistency
- Strong Guarantees: Clear behavioral expectations during failures
- ACID Support: Full transactional capabilities
Challenges of CP Systems
- Reduced Availability: System becomes unavailable during partitions
- Higher Latency: Consensus protocols add latency to operations
- Scaling Complexity: More complex to scale than AP systems
- Single Points of Coordination: Consensus can become a bottleneck
- Geographic Limitations: Cross-region consistency can be slow
When to Choose CP Systems
Ideal Scenarios:
- Financial and banking applications
- Legal and compliance systems
- Healthcare record management
- Inventory with strict accuracy requirements
- Any system where incorrect data is worse than temporary unavailability
Avoid When:
- User experience is more important than perfect consistency
- High write volumes with relaxed consistency requirements
- Global applications where partition tolerance is frequent
- Systems requiring sub-millisecond response times
Best Practices for CP System Design
- Design for Partitions: Plan explicitly for partition scenarios
- Monitor Quorum Health: Implement robust monitoring and alerting
- Optimize for Common Case: Design for normal operations while handling edge cases
- Clear Failure Modes: Make system behavior predictable during failures
- Regular Testing: Use chaos engineering to test partition scenarios
- Geographic Considerations: Understand latency implications of global consistency
Conclusion
CP systems are essential for applications where data accuracy and integrity cannot be compromised, even temporarily. While they sacrifice availability during network partitions, they provide the strong consistency guarantees required for critical business operations.
The choice of a CP system should be deliberate, based on careful analysis of business requirements, regulatory needs, and user expectations. When implemented correctly, CP systems provide the reliable foundation necessary for mission-critical applications where “approximately correct” is not acceptable.
AP Systems: Prioritizing Availability Over Consistency During Partitions
Understanding AP Systems
For a distributed system that must handle network partitions, an AP system chooses to prioritize Availability (A) over Consistency (C) during such an event. This means that if a network partition occurs, the system will continue to operate and respond to requests within each isolated partition. It will allow reads and writes to continue, even if it cannot immediately guarantee that all nodes have the exact same, most up-to-date data.
The core principle here is that users should always be able to interact with the system, even if the data they see might be temporarily stale or inconsistent. When the network partition heals, the system will then work to resolve any inconsistencies that arose, eventually bringing all nodes back to a consistent state (this is known as eventual consistency). This model relies on “optimistic concurrency,” assuming that conflicts will be rare or can be resolved later.
Key Characteristics of AP Systems
- Continuous Operation: Services remain operational during network partitions
- Eventual Consistency: Data converges to consistency over time
- Optimistic Concurrency: Assumes conflicts are rare and resolvable
- Partition Independence: Each partition can operate independently
- Conflict Resolution: Mechanisms to handle divergent data when partitions heal
- High Availability: Designed for maximum uptime and responsiveness
Real-World Examples of AP Systems in Action
Social Media Feeds and Messaging
Scenario: When you open your social media app, you expect to see your feed or messages immediately.
It’s far more acceptable to see a post from 30 seconds ago, or for a message to arrive a few seconds late, than to see a “service unavailable” error. If a network partition occurs, different parts of the social media platform (e.g., user profiles in one data center, feed items in another) will continue to operate.
Partition Behavior: You might see a slightly older version of a friend’s profile or miss a very recent post, but you can still browse, post, and interact. The system will sync up all the data later.
Why AP is Optimal:
- User engagement is more important than perfect consistency
- Social interactions can tolerate slight delays
- Massive user bases require continuous availability
- Network effects depend on uninterrupted access
Trade-off: Temporary inconsistency for continuous social interaction
Online Shopping Carts (Initial Stages)
Scenario: Adding items to your shopping cart on an e-commerce site.
The system typically prioritizes availability for the browsing and cart experience. If a network partition temporarily prevents your session’s server from communicating with the absolute latest inventory numbers, it will still allow you to add items. The cart will remain available and responsive.
Partition Behavior: The true inventory check and strong consistency typically only happen at the checkout stage (a critical transaction point), where ACID-like guarantees are needed. Until then, keeping the cart experience fluid and available is key.
Why AP is Optimal:
- Shopping experience should be seamless
- Cart abandonment increases with any friction
- Inventory validation can happen at checkout
- User experience drives conversion rates
IoT Data Ingestion and Telemetry
Scenario: Large-scale Internet of Things (IoT) deployments generating vast amounts of sensor data.
These systems must be highly available to continuously ingest data, even if network connectivity between sensors and central processing units is intermittent or partitioned. It’s more important to record all temperature readings or device statuses than to ensure every single reading is immediately propagated to all analytics dashboards globally.
Partition Behavior: Data collection continues in all partitions, with data eventually converging when connectivity is restored. Analysis can be performed on the complete dataset later.
Why AP is Essential:
- Sensor data is time-sensitive and cannot be re-collected
- Analytics can tolerate slight delays
- Massive scale requires continuous ingestion
- Data loss is worse than temporary inconsistency
Real-time Gaming Leaderboards
Scenario: Leaderboards for casual games.
It’s generally more important that the leaderboard loads quickly and is always visible, rather than being absolutely perfectly updated in real-time down to the millisecond. If a network partition means a player’s latest score takes a few seconds to appear for others, that’s acceptable.
Partition Behavior: The system remains available, allowing players to check their rankings without interruption, even if the data is eventually consistent.
Why AP Works:
- Player engagement depends on responsive interfaces
- Slight ranking delays don’t affect gameplay
- Competitive integrity can be maintained through eventual consistency
- Social gaming features require continuous availability
Content Delivery Networks (CDNs)
Scenario: Delivering static and dynamic content globally.
CDNs are designed to deliver content with extremely high availability and low latency. If a specific edge server or a region’s data center becomes partitioned from the main origin server, the CDN nodes in that partition will continue to serve cached content.
Partition Behavior: While this content might not be the absolute latest version (e.g., a newly updated image might not be propagated yet), it ensures that users can still access the website or media without interruption.
Why AP is Critical:
- User expectations for instant content loading
- Content freshness is less critical than availability
- Global distribution inherently creates partition scenarios
- Revenue depends on continuous content delivery
Collaborative Document Editing
Scenario: Multiple users editing shared documents (like Google Docs in offline mode).
Partition Behavior: When users are offline or partitioned, they can continue editing their local copy. When connectivity is restored, the system merges changes using operational transformation or conflict-free replicated data types (CRDTs).
Why AP Enables Productivity:
- Users can’t wait for perfect connectivity to work
- Productivity requires continuous access to documents
- Conflict resolution algorithms handle most merge scenarios
- Collaboration benefits from optimistic concurrency
Database Choices for AP Systems
Primary Recommendation: Apache Cassandra
Why Cassandra is ideal for AP systems:
- Masterless Architecture: Every node can accept writes and serve reads independently
- Tunable Consistency: Configurable consistency levels from “ANY” (highest availability) to “ALL” (highest consistency)
- Partition Resilience: Continues operating when individual nodes or data centers fail
- Eventual Consistency: Built-in mechanisms for data convergence
- Linear Scalability: Performance scales linearly with additional nodes
- Multi-Data Center Support: Designed for geographic distribution
Key AP Features:
- Hinted Handoffs: Stores writes for temporarily unavailable nodes
- Read Repair: Fixes inconsistencies during read operations
- Anti-Entropy: Background processes ensure eventual consistency
- Merkle Trees: Efficient detection of inconsistent data
- Gossip Protocol: Decentralized cluster state management
Consistency Levels for High Availability:
- ANY: Write succeeds when at least one node acknowledges
- ONE: Read/write from a single node
- QUORUM: Majority of nodes (balanced approach)
Alternative AP Database Options
Amazon DynamoDB:
- Fully managed NoSQL with automatic scaling
- Eventually consistent reads by default
- Global tables for multi-region replication
- Built-in partition tolerance and high availability
MongoDB (with specific configuration):
- Can be configured for high availability over consistency
- Replica sets with read preferences
- Sharding for horizontal scaling
- Flexible document model
Apache CouchDB:
- Multi-version concurrency control
- Bi-directional replication
- Conflict detection and resolution
- Designed for offline-first applications
Redis Cluster:
- In-memory data structure store
- Automatic partitioning
- Continues operating during node failures
- Excellent for caching and session storage
Riak:
- Distributed key-value store
- Configurable N/R/W values
- Built-in conflict resolution
- Designed for high availability
Implementation Strategies for AP Systems
Conflict Resolution Mechanisms
Last Writer Wins (LWW):
- Simple timestamp-based resolution
- Works well for low-conflict scenarios
- May lose data in high-conflict situations
Vector Clocks:
- Tracks causality between updates
- Enables more sophisticated conflict detection
- Used by systems like Riak and Voldemort
Conflict-Free Replicated Data Types (CRDTs):
- Mathematically proven to converge
- No conflicts by design
- Ideal for collaborative applications
Application-Level Resolution:
- Custom business logic for conflict handling
- Most flexible but requires careful design
- Can incorporate domain-specific rules
Data Modeling for AP Systems
Denormalization:
- Duplicate data to reduce cross-partition dependencies
- Optimize for read patterns
- Accept storage overhead for availability
Event Sourcing:
- Store events rather than current state
- Natural fit for eventual consistency
- Enables time-travel and audit capabilities
Saga Pattern:
- Manage distributed transactions
- Compensating actions for rollback
- Maintains availability during long-running processes
Monitoring and Observability
Key Metrics for AP Systems
Availability Metrics:
- Uptime and service level objectives (SLOs)
- Response time percentiles
- Error rates and types
Consistency Metrics:
- Replica lag and convergence time
- Conflict detection and resolution rates
- Data integrity checks
Partition Metrics:
- Network partition detection
- Partition duration and frequency
- Cross-partition communication patterns
Alerting Strategies
Availability Alerts:
- Service unavailability
- High error rates
- Response time degradation
Consistency Alerts:
- Excessive replica lag
- High conflict rates
- Data integrity violations
Trade-offs and Considerations
Benefits of AP Systems
- High Availability: Continuous operation during failures
- Scalability: Linear scaling with additional nodes
- Global Distribution: Works well across regions
- User Experience: Responsive interfaces and interactions
- Fault Tolerance: Graceful degradation during failures
Challenges of AP Systems
- Eventual Consistency: Temporary data inconsistencies
- Conflict Resolution: Complex handling of divergent updates
- Data Modeling: Requires different approaches than traditional RDBMS
- Debugging Complexity: More complex failure modes
- Business Logic: Applications must handle inconsistency
When to Choose AP Systems
Ideal Scenarios:
- Social media and content platforms
- IoT and sensor data collection
- Content delivery and caching
- Collaborative applications
- High-volume, low-latency services
- Global applications with geographic distribution
Avoid When:
- Financial transactions requiring immediate consistency
- Systems where data accuracy is more important than availability
- Applications with complex transactional requirements
- Scenarios where conflict resolution is impossible
Best Practices for AP System Design
- Design for Eventual Consistency: Build applications that can handle temporary inconsistencies
- Implement Robust Conflict Resolution: Choose appropriate strategies for your use case
- Monitor Consistency Lag: Track how quickly data converges
- Test Partition Scenarios: Use chaos engineering to validate behavior
- Optimize for Common Patterns: Design data models for typical access patterns
- Plan for Conflict Resolution: Have clear strategies for handling divergent data
Conclusion
AP systems are essential for modern applications that prioritize user experience, global scale, and continuous operation. While they require careful consideration of eventual consistency and conflict resolution, they enable the responsive, always-available services that users expect in today’s connected world.
The choice of an AP system should be based on understanding that temporary inconsistencies are acceptable trade-offs for continuous availability. When designed correctly, AP systems provide the foundation for scalable, resilient applications that can handle the demands of modern distributed computing while maintaining excellent user experiences.
A short, opinionated path through the CAP literature, with citations rendered from the site’s bibliography.bib. Read in this order; each paper corrects a misunderstanding the previous one tends to leave behind.
The one-line version
CAP is not a menu of three where you pick two. Partitions happen whether you like it or not; the theorem is about what your system does while one is happening.
1. The conjecture
Brewer presented the idea as a keynote, not a proof, and framed it in terms of the trade-offs web services were already making in practice (Brewer, 2000). The original slides are informal; what later became “the CAP theorem” was a slogan first.
2. The proof, and the narrowing
Gilbert and Lynch turned the conjecture into a theorem by pinning down what “consistent” and “available” mean, and that is where the formal result becomes narrow: linearizable consistency plus every request receiving a response, in an asynchronous network (Gilbert & Lynch, 2002). Most production systems do not aim for either definition in full.
3. Twelve years later
Brewer revisited the theorem to push back on the “two out of three” reading. Partitions are rare, and systems can be consistent and available almost all of the time; the design question is detection, a sensible degraded mode, and recovery (E. Brewer, 2012).
4. The trade-off that applies even without partitions
Abadi’s PACELC adds the missing axis: when the network is fine, replicated systems still trade latency against consistency. Many “AP” databases are really “EL” ones that chose low latency (Abadi, 2012).
5. What you can keep under partition
Bailis and colleagues catalogue which transactional guarantees survive high availability and which cannot, a useful map for anyone choosing isolation levels in a distributed database (Bailis et al., 2013).
6. Why the vocabulary fails
Kleppmann’s critique argues that CAP’s definitions are too blunt for real engineering discussions and proposes talking about delay-sensitivity instead. If you read only one of these, read this one (Kleppmann, 2015).
7. How CP systems get consistency in practice
Raft is the consensus algorithm most of the CP databases on this site rely on. It was designed explicitly to be understandable, and the paper delivers on that (Ongaro & Ousterhout, 2014).
Where this connects on the site
- System Design/partition_tolerance_in_distributed_systems cites the same sources inline.
- System Design/cp_systems_design and System Design/ap_systems_design show the two choices under partition.
Reading order: conjecture → proof → twelve years later → PACELC → HAT → critique
An exhaustive, alphabetically sorted list of software architecture concepts, their creators, brief descriptions, and links to key resources, preferably by the authors.
A
ACID Properties - Theo Härder and Andreas Reuter
- Description: A set of properties that guarantee reliable processing in database transactions: Atomicity, Consistency, Isolation, and Durability.
- Link: Principles of Transaction-Oriented Database Recovery (Original paper)
Active Record Pattern - Martin Fowler
- Description: An architectural pattern for accessing data in a database, where each object wraps a row in a database table or view.
- Link: Active Record (Article by Martin Fowler)
Additive Programming - Chris Hanson and Gerald Jay Sussman
- Description: A programming style where functionality is added incrementally, allowing for flexible and extensible software design without modifying existing code.
- Link: Software Design for Flexibility (Book by Hanson and Sussman)
Adapter Pattern - Gang of Four (Erich Gamma, Richard Helm, Ralph Johnson, John Vlissides)
- Description: Allows incompatible interfaces to work together by wrapping an existing class with a new interface.
- Link: Design Patterns: Elements of Reusable Object-Oriented Software (Book by the Gang of Four)
Actor Model - Carl Hewitt, Peter Bishop, Richard Steiger
- Description: A mathematical model of concurrent computation that treats “actors” as the universal primitives of concurrent digital computation.
- Link: A Universal Modular ACTOR Formalism (Original paper)
Anti-Corruption Layer (from DDD) - Eric Evans
- Description: A layer that isolates a domain model from external systems to prevent their influence on the domain design, translating between the two as necessary.
- Link: Domain-Driven Design Reference (Official reference by Eric Evans)
Aspect-Oriented Programming (AOP) - Gregor Kiczales
- Description: A programming paradigm that increases modularity by allowing the separation of cross-cutting concerns, such as logging or security.
- Link: Aspect-Oriented Programming (Original paper)
How this page is organized
One entry per concept: who introduced or popularized it, a one-paragraph description, and a primary source. Entries are alphabetical, so use the table of contents on the right or the search box rather than scrolling.
Reading suggestions (click to expand)
- New to architecture: start with Coupling, Cohesion, Modularity and Separation of Concerns, then Hexagonal Architecture.
- Preparing for interviews: pair this glossary with the trade-offs note and the learning paths.
- Everything here links to a primary source; when two sources disagree, the original author’s definition wins.
B
BASE Properties - Eric Brewer (coined term), Dan Pritchett (popularized)
- Description: An alternative to ACID properties, BASE stands for Basically Available, Soft state, Eventual consistency. It provides a model for designing distributed systems that prefer availability over strong consistency.
- Link: BASE: An ACID Alternative (Article by Dan Pritchett)
Backend for Frontend (BFF) - Sam Newman
- Description: Creates separate backend services for specific frontend applications or interfaces, optimizing each for its needs.
- Link: Pattern: Backends For Frontends (Article by Sam Newman)
Big Ball of Mud - Brian Foote and Joseph Yoder
- Description: A software system that lacks a perceivable architecture; it’s a haphazardly structured, sprawling, sloppy, spaghetti-code jungle.
- Link: Big Ball of Mud (Original paper)
Bridge Pattern - Gang of Four
- Description: Decouples an abstraction from its implementation so that the two can vary independently.
- Link: Design Patterns Book (Book by the Gang of Four)
Bulkhead Pattern - Michael Nygard
- Description: Isolates components or services to prevent cascading failures, improving system resilience.
- Link: Release It! (Book by Michael Nygard)
C
C4 Model - Simon Brown
- Description: A simple hierarchical way to visualize software architecture using a set of diagrams: Context, Container, Component, and Code.
- Link: The C4 Model for Software Architecture (Official website by Simon Brown)
CAP Theorem - Eric Brewer
- Description: States that in the presence of a network partition, a distributed system can provide either Consistency or Availability, but not both.
- Link: CAP Twelve Years Later: How the “Rules” Have Changed (Article by Eric Brewer)
Circuit Breaker Pattern - Michael Nygard
- Description: Detects failures and encapsulates logic to prevent cascading failures, enhancing system stability.
- Link: Release It! (Book by Michael Nygard)
Clean Architecture - Robert C. Martin (Uncle Bob)
- Description: Separates the elements of design into ringed layers with dependencies pointing inwards, promoting separation of concerns.
- Link: The Clean Architecture (Blog post by Uncle Bob)
Cleanroom Software Engineering - Harlan Mills
- Description: A software development process intended to produce software with a certifiable level of reliability, using formal methods and statistical quality control.
- Link: Cleanroom Software Engineering (Paper by Harlan Mills)
Client-Server Architecture - Evolved Concept
- Description: A network architecture where client devices request resources and services from centralized servers.
- Link: Client-Server Model (Explanation)
Command Pattern - Gang of Four
- Description: Encapsulates a request as an object, allowing for parameterization and queuing of requests.
- Link: Design Patterns Book
Command Query Responsibility Segregation (CQRS) - Greg Young
- Description: Segregates the read and write operations, using separate models optimized for each, improving scalability.
- Link: CQRS Documents (Greg Young’s blog)
Communicating Sequential Processes (CSP) - Tony Hoare
- Description: A formal language for describing patterns of interaction in concurrent systems through message-passing communication.
- Link: Communicating Sequential Processes (Book by Tony Hoare)
Constraint Satisfaction Problems (CSP) - Evolved Concept
- Description: Mathematical problems defined as a set of objects whose state must satisfy a number of constraints or limitations.
- Link: Constraint Satisfaction Problems (Explanation)
D
Data Mapper Pattern - Martin Fowler
- Description: Separates the in-memory objects from the database, mapping between the two, allowing both to vary independently.
- Link: Data Mapper (Article by Martin Fowler)
Data Transfer Object (DTO) - Martin Fowler
- Description: An object that carries data between processes to reduce the number of method calls, often used in remote interfaces.
- Link: Data Transfer Object (Article by Martin Fowler)
Design by Contract - Bertrand Meyer
- Description: Designing software by defining formal, precise, and verifiable interface specifications for software components.
- Link: Object-Oriented Software Construction (Book by Bertrand Meyer)
Dependency Injection (DI) - Martin Fowler (popularized the term)
- Description: A technique where an object receives other objects it depends on, called dependencies, from an external source rather than creating them itself.
- Link: Inversion of Control Containers and the Dependency Injection pattern (Article by Martin Fowler)
Dependency Inversion Principle - Robert C. Martin
- Description: States that high-level modules should not depend on low-level modules; both should depend on abstractions.
- Link: The Dependency Inversion Principle (Archived article by Uncle Bob)
Design Patterns (Gang of Four Patterns) - Erich Gamma, Richard Helm, Ralph Johnson, John Vlissides
- Description: A catalog of common solutions to recurring design problems in software development.
- Link: Design Patterns: Elements of Reusable Object-Oriented Software (Book by the Gang of Four)
Domain-Driven Design (DDD) - Eric Evans
- Description: An approach to software development that focuses on modeling software to match a domain according to input from domain experts.
- Link: Domain-Driven Design Reference (Official reference by Eric Evans)
Domain-Specific Languages (DSLs) - Chris Hanson and Gerald Jay Sussman
- Description: Specialized computer languages focused on a particular aspect of a software system, designed to be expressive in that domain.
- Link: Software Design for Flexibility (Book)
DRY Principle (Don’t Repeat Yourself) - Andy Hunt and Dave Thomas
- Description: A principle aiming to reduce repetition of software patterns, replacing them with abstractions or data normalization.
- Link: The Pragmatic Programmer (Book by Hunt and Thomas)
E
Enterprise Integration Patterns - Gregor Hohpe and Bobby Woolf
- Description: A catalog of patterns for designing messaging systems within enterprise applications to facilitate integration.
- Link: Enterprise Integration Patterns (Official website)
Event Sourcing - Martin Fowler and Greg Young
- Description: Stores state changes as a sequence of events, allowing for full audit trails and state reconstruction.
- Link: Event Sourcing (Article by Martin Fowler)
Event Storming - Alberto Brandolini
- Description: A workshop-based method to quickly model complex business domains by focusing on domain events, facilitating collaboration.
- Link: Event Storming (Official website)
Event-Driven Architecture (EDA) - Evolved Concept
- Description: Promotes the production, detection, consumption of, and reaction to events, enabling loose coupling.
- Link: Event-Driven Architecture (Article by Martin Fowler)
Eventual Consistency - Werner Vogels
- Description: A consistency model used in distributed computing to achieve high availability by allowing data to be temporarily inconsistent.
- Link: Eventually Consistent (Article by Werner Vogels)
F
Formal Methods in Software Engineering - Evolved Concept
- Description: Mathematical techniques and tools for specifying, developing, and verifying software and hardware systems to ensure correctness and reliability.
- Link: Formal Methods Wiki (General resource)
Functional Programming Concepts
Currying - Haskell Curry
- Description: Transforming a function that takes multiple arguments into a sequence of functions each taking a single argument.
- Link: Currying in Functional Programming (Paper)
Monads - Philip Wadler (popularized in programming)
- Description: A design pattern used to handle program side effects in functional programming, encapsulating values along with a context.
- Link: Monads for Functional Programming (Paper by Philip Wadler)
Higher-Order Functions - Evolved Concept
- Description: Functions that take other functions as arguments or return them as results, fundamental in functional programming languages.
- Link: Higher-Order Functions (Explanation)
Immutability - Evolved Concept
- Description: Data not changing state after creation, enhancing predictability and thread safety in functional programming.
- Link: Immutability in Functional Programming (Article)
Lazy Evaluation - Haskell Community
- Description: An evaluation strategy which delays the evaluation of an expression until its value is needed.
- Link: Lazy Evaluation (Explanation)
H
Hillel Wayne’s Contributions to Formal Methods
- Description: Hillel Wayne advocates for the practical application of formal methods in software engineering, making formal verification techniques accessible to industry practitioners.
- Link: Hillel Wayne’s Website (Official website)
Hoare Logic - C.A.R. Hoare
- Description: A formal system used to reason about the correctness of computer programs, using preconditions and postconditions.
- Link: An Axiomatic Basis for Computer Programming (Original paper by C.A.R. Hoare)
Horn Clauses - Alain Colmerauer and Philippe Roussel
- Description: A special kind of clause used in logic programming, especially in Prolog, representing implications suitable for computation.
- Link: The Birth of Prolog (Historical paper)
Hexagonal Architecture (Ports and Adapters) - Alistair Cockburn
- Description: Promotes designing software applications to be loosely coupled and easily testable by isolating the core logic from external factors.
- Link: Hexagonal Architecture (Official website by Alistair Cockburn)
I
Immutable Infrastructure - Kief Morris
- Description: An approach where servers or systems are never modified after deployment, ensuring consistency across environments and simplifying deployment.
- Link: Infrastructure as Code (Book by Kief Morris)
Inversion of Control (IoC) - Michael Mattsson (coined term), Martin Fowler (popularized)
- Description: A design principle in which custom-written portions of a program receive the flow of control from a generic framework.
- Link: Inversion of Control Containers and the Dependency Injection pattern (Article by Martin Fowler)
K
KISS Principle (Keep It Simple, Stupid) - Kelly Johnson (Originated in U.S. Navy)
- Description: States that simplicity should be a key goal in design and unnecessary complexity should be avoided.
- Link: KISS Principle (Background information)
L
Lambda Calculus - Alonzo Church
- Description: A formal system in mathematical logic for expressing computation based on function abstraction and application.
- Link: An Unsolvable Problem of Elementary Number Theory (Original paper)
Layered Architecture (4+1 View Model) - Philippe Kruchten
- Description: Describes a software architecture using five concurrent views, addressing different concerns for various stakeholders.
- Link: Architectural Blueprints—The “4+1” View Model of Software Architecture (Paper by Philippe Kruchten)
Layered Architecture (N-tier Architecture) - Evolved Concept
- Description: Organizes software systems into layers, each with a specific role, promoting separation of concerns.
- Link: Patterns of Enterprise Application Architecture (Book by Martin Fowler)
Layered Architecture (Onion Architecture) - Jeffrey Palermo
- Description: Places the domain model at the core, with layers wrapping around it, aiming for separation of concerns.
- Link: The Onion Architecture (Blog series by Jeffrey Palermo)
Logic Programming Concepts
Unification - J. Alan Robinson
- Description: An operation that makes two logical expressions identical by finding a substitution, fundamental in logic programming.
- Link: A Machine-Oriented Logic Based on the Resolution Principle (Original paper)
Resolution Principle - J. Alan Robinson
- Description: A rule of inference leading to a refutation-complete theorem-proving technique for first-order logic.
- Link: [Same as above]
SLD Resolution - Robert Kowalski
- Description: A refinement of resolution used in logic programming to derive consequences from a set of clauses.
- Link: Logic for Problem Solving (Book by Robert Kowalski)
M
Micro Frontends - Luca Mezzalira
- Description: Extends microservices to frontend development, allowing teams to independently develop and deploy UI components.
- Link: Micro Frontends (Official website)
Microkernel Architecture - Evolved Concept
- Description: Separates core functionality from extended functionality, promoting plug-in modules for extensibility.
- Link: Microkernel Architecture Pattern (Article by Mark Richards)
Microservices Architecture - James Lewis and Martin Fowler
- Description: Structures an application as a collection of loosely coupled services, which implement business capabilities.
- Link: Microservices (Article by Martin Fowler)
Model-Driven Architecture (MDA) - Object Management Group (OMG)
- Description: Relies on modeling and model transformations to define the functionality and behavior of a system.
- Link: Model-Driven Architecture (Official website)
Model Checking - E. Clarke, E. Emerson, and A. Pnueli
- Description: An automated technique that, given a finite-state model of a system and a formal property, systematically checks whether this property holds for that model.
- Link: Model Checking (Book by Clarke, Grumberg, and Peled)
Model-View-Controller (MVC) - Trygve Reenskaug
- Description: Separates an application into three main components: Model (data), View (UI), and Controller (business logic).
- Link: MVC - XEROX PARC 1978-79 (Original reports by Reenskaug)
Model-View-Presenter (MVP) - Mike Potel
- Description: Improves the separation of concerns between the user interface and the business logic by introducing a Presenter.
- Link: MVP Design Pattern (Original paper by Mike Potel)
Model-View-ViewModel (MVVM) - John Gossman
- Description: Facilitates the separation of the development of the graphical user interface from the business logic or back-end logic.
- Link: Introduction to Model/View/ViewModel pattern (Archived blog post by John Gossman)
Monitoring Concepts
Observability - Charity Majors and Others
- Description: A measure of how well internal states of a system can be inferred from its external outputs, crucial for debugging and performance.
- Link: Observability—A 3-Year Retrospective (Article by Charity Majors)
The RED Method - Tom Wilkie
- Description: A monitoring approach focused on three key metrics: Rate, Errors, and Duration for microservices.
- Link: The RED Method (Article by Tom Wilkie)
The USE Method - Brendan Gregg
- Description: Analyzing system performance by checking Utilization, Saturation, and Errors of resources.
- Link: The USE Method (Article by Brendan Gregg)
Service Level Objectives (SLOs) - Google SRE Team
- Description: A target level of reliability for a service, defined to ensure acceptable performance and availability.
- Link: SRE Book: Service Level Objectives (Chapter from Google’s SRE Book)
Blackbox and Whitebox Monitoring - Evolved Concepts
- Description: Blackbox monitoring treats the system as a whole without internal insight, while whitebox monitoring uses internal metrics.
- Link: Blackbox vs Whitebox Monitoring (Article)
S
SMT Solvers (Satisfiability Modulo Theories) - Evolved Concept
- Description: Tools that decide the satisfiability of logical formulas with respect to combinations of background theories, used in proof automation.
- Link: Introduction to SMT (Book)
Software Specifications - Evolved Concept
- Description: Detailed descriptions of software system requirements, including functional and non-functional aspects, serving as a blueprint for development.
- Link: Writing Good Software Specifications (Article by Hillel Wayne)
T
TLA+ (Temporal Logic of Actions) - Leslie Lamport
- Description: A formal specification language developed for designing, modeling, documenting, and verifying concurrent and distributed systems.
- Link: TLA+ Homepage (Official website by Leslie Lamport)
Test-Driven Development (TDD) - Kent Beck
- Description: A software development process that relies on the repetition of a very short development cycle: write a failing test, make it pass, and refactor.
- Link: Test-Driven Development: By Example (Book by Kent Beck)
The Twelve-Factor App - Adam Wiggins
- Description: A methodology for building software-as-a-service apps that use declarative formats for setup automation, aiming for portability and resilience.
- Link: The Twelve-Factor App (Official website by Adam Wiggins)
Why determinism is the most underrated principle in software engineering — and how to apply it at every level
The Boring Reason Systems Fail
Most software systems don’t fail because of microservices. They don’t fail because of monoliths. They don’t fail because of agile. They fail for a much more boring reason: they’re unpredictable.
If you make a change and can’t reliably determine the impact of that change, you can’t safely evolve your system. And if you can’t evolve your system, it’s already a legacy system — regardless of when it was written.
This is a problem of determinism: the property that the same input always produces the same output. It sounds obvious. It’s anything but.
Determinism is the bridge between “it works on my machine” and “it works every time, everywhere.” And if you care about continuous delivery, about evolutionary architecture, about building systems that improve instead of decay — you should care deeply about determinism.
The Chain: Why Determinism Enables Everything Else
There’s a logical chain that connects determinism to your ability to evolve a system:

You can’t evolve what you can’t measure. You can’t measure what you can’t repeat. And you can’t repeat something that isn’t deterministic.
If your tests are flaky, if your builds sometimes fail “for reasons,” if concurrency randomly breaks things — then your delivery pipeline stops being a real learning system. It becomes release theater: going through the motions without actually validating anything.
Determinism is the prerequisite for trust. Without it, your tests lie to you. Your pipeline lies to you. Your architecture rots quietly in the background.
With it, you can run thousands of experiments per day. You can detect unintended consequences. You can move fast without gambling. That’s an evolutionary capability.
The Four Enemies of Determinism
Non-determinism doesn’t come from one place. It sneaks in through at least four distinct channels, each requiring a different fix.

Let’s examine each one.
Enemy 1: Uncontrolled Time
If your code calls datetime.now() directly, you’ve just injected non-determinism into your system. The same input tomorrow produces a different output. That’s not testable. That’s not reproducible. That’s not evolutionary.
This is one of the most common sources of non-determinism, and one of the easiest to fix.

The Fix: Pass Time as Data
Instead of reaching for the system clock inside your business logic, inject a clock interface. Treat “now” as an input parameter, not a global variable.
Something magical happens when you do this:
- Freeze time — test what happens at exactly midnight on December 31st
- Fast-forward — simulate a 30-day expiry window in milliseconds
- Replay production bugs — feed in the exact timestamp from the production log
- Test DST transitions — create a clock that jumps forward/backward at will
You’ve turned the universe into a parameter. That’s a powerful tool.
Real-World Example
In our codebase, our OXI outage detection system runs differently depending on business hours (6 AM–11 PM in a specific timezone). Without clock injection, testing this logic requires either running tests at specific times of day or hacking the system clock. With clock injection, we test every hour of every timezone in milliseconds.
Enemy 2: Shared Mutable State
Concurrency is where determinism goes to die.
If thread scheduling decides the order of execution, your architecture is now probabilistic. It’s subject to random race conditions and Heisenbugs — bugs that disappear when you try to observe them. You’ll be familiar with the phrase “well, it works 99% of the time.” That’s not engineering. That’s roulette.
The Fix: Remove the Mutation
There are effective ways to handle concurrency without sacrificing determinism:
- Actor models — each actor processes messages sequentially, no shared state
- Single-threaded event loops — one thread, deterministic ordering
- Immutable state — when state doesn’t mutate unpredictably, order stops mattering quite so much
- Deterministic merge strategies — when you must combine results, make the merge itself deterministic
The key insight: when execution is structured as input → decision → event, you regain control. Each step is a pure transformation that can be tested independently.
What This Looks Like at Scale
Dave Farley describes a system at LMAX that processed global financial trades where each service was completely deterministic. Given the same starting state and the same sequence of events, they got exactly the same result every time. This isn’t academic purity — it’s how you build systems where evolutionary change is safe, because you can always verify the impact.
The Central Pattern: Deterministic Core + Imperative Shell
If you take one idea away from this entire article, take this one: separate the code that decides from the code that acts.

The Deterministic Core
This is where your business logic lives. Pure functions that take state in and produce decisions out. No database interactions. No clock. No randomness. No network. No side effects.
Because the core is pure:
- You don’t need mocks
- You don’t need frameworks
- You don’t need complex test scaffolding
- You pass in state, you assert on output
- You can run thousands of tests in milliseconds
Those are your fitness functions at scale. And now your architecture can evolve safely, one bit at a time.
The Imperative Shell
This is where you deal with the messy world: talk to the database, talk to the network, read the clock, execute side effects. The shell is thin, mechanical, and boring by design. It converts the outside world into inputs for the core, and converts the core’s decisions into actions.
Why This Works
This pattern isn’t new. It’s separation of concerns. It’s hexagonal architecture. It’s ports and adapters. But framing it through the lens of determinism makes the benefit crystal clear: the core is where your tests give you confidence, and the shell is where you manage the chaos.
When the core is pure, testing is trivially fast. When the shell is thin, there’s less surface area for non-determinism to hide in. The result is a system you can change with confidence.
Enemy 3: Uncontrolled State Space
Non-determinism isn’t only about threads or clocks. It’s about uncontrolled state space.
If your system has 500 possible implicit states and you don’t know which one you’re in at any given moment, you don’t have architecture — you have entropy.

The Fix: Make State Explicit
Design narrower scopes. Use explicit state machines. Make boundaries between components serving different purposes clear. Design components to reject invalid states.
The narrower the scope, the more control you have over state. The more control you have over state, the more deterministic the system becomes. And the more deterministic it becomes, the more confidently you can change it.
What Explicit State Looks Like
Instead of a boolean isActive and a nullable completedAt and a string status that might be any of 12 values — use a state machine with 4 defined states and 5 valid transitions. Now:
- You always know which state you’re in
- Invalid transitions are rejected at compile time or runtime
- Every state transition is testable
- The system’s behavior is predictable at every step
This is evolutionary architecture in practice: not a grand upfront design, but structural choices that keep the system changeable.
Enemy 4: Environment Coupling
Determinism doesn’t stop at code. It extends to your entire software supply chain.

Hermetic Builds
If the same source code produces different artifacts depending on when or where you build it, your build is non-deterministic. Hermetic builds — builds that are fully self-contained and reproducible — ensure that the same source always produces the same binary.
Pinned Dependencies
If your build pulls “latest” versions of dependencies, you’ve introduced non-determinism at the supply chain level. A library update on Tuesday could break your Friday build, and you’d have no idea what changed. Pin your dependencies. Lock your versions.
Idempotent Deployments
If applying your deployment twice changes the result, your infrastructure is non-deterministic and your production environment is, as Dave Farley puts it, “a mystery wrapped in an enigma.”
Idempotent deployments mean: run it once, run it ten times — same result. Infrastructure as code. Declarative configuration. No manual steps that someone might forget or do differently.
The Payoff: Speed of Learning
Most organizations optimize for features. But the best organizations optimize for speed of learning.

Deterministic systems deliver three compounding benefits:
1. Shorter Feedback Loops
When tests are deterministic, you know immediately whether a change broke something. No re-runs “just in case.” No “it was probably a flake.” The feedback is instant and trustworthy.
2. Reduced Cognitive Load
When the system is predictable, developers don’t need to hold the entire state space in their heads. They can reason locally about the component they’re changing, because the boundaries are clear and the behavior is consistent.
3. Reproducible Debugging
When a bug happens in production, you can reproduce it locally by feeding in the same inputs. No more “I can’t reproduce it” — because the system is deterministic, the same inputs always produce the same failure.
These benefits compound. Shorter loops mean more experiments. More experiments mean faster learning. Faster learning means better software. This is why continuous delivery works — not because of pipelines, but because of determinism. The pipeline is just the amplifier that helps you see how close you are to it.
Practical Checklist
Here’s how to increase determinism in your system, starting today:
Code Level
- Inject clocks instead of calling
now()directly - Separate pure business logic from side-effect code (core + shell)
- Use immutable data structures by default
- Avoid shared mutable state between threads
- Make random number generators injectable/seedable
Testing Level
- Eliminate flaky tests — each one is a determinism leak
- Use deterministic assertions (exact values, not “not null”)
- Make test data factories produce reproducible output
- Run tests in isolation — no shared database state between tests
- Treat a flaky test as a P1 bug, not an annoyance
Build Level
- Pin all dependency versions (lock files, exact versions)
- Use hermetic builds (same source = same artifact, regardless of when/where)
- Cache build outputs by content hash, not by timestamp
- Verify builds are reproducible by building twice and comparing
Deployment Level
- Make deployments idempotent (apply twice = same result)
- Use infrastructure as code with version-controlled configuration
- Eliminate manual deployment steps
- Test rollbacks — they should produce a known prior state
The Key Insight
Determinism isn’t a coding trick. It’s a systems property. It’s not about purity for its own sake — it’s about building systems where change is safe, feedback is trustworthy, and evolution is possible.
“Increasing determinism is the target. The pipeline is really just the amplifier that helps us to see how close we are to it.”
Every time you inject a clock instead of calling now(), every time you extract a pure function from a side-effecting method, every time you pin a dependency or make a deployment idempotent — you’re making your system more deterministic. And a more deterministic system is one that can evolve.
That’s the simplest way to make your architecture testable and reproducible. It works every time.
System Design Learning Paths: Choose Your Journey to Mastery
System design is a vast domain spanning theoretical foundations, practical implementation, and strategic thinking. With the wealth of content available, the biggest challenge isn’t finding resources—it’s knowing where to start and how to progress systematically toward your goals.
This guide provides structured learning paths tailored to different backgrounds, timelines, and objectives. Whether you’re preparing for interviews, transitioning to senior roles, or building deep distributed systems expertise, there’s a path designed for your journey.
Track your progress (click to expand)
Copy this list into your own notes and tick items as you go. Obsidian’s custom task markers survive the build, so
>marks work in progress and?marks something to revisit.
- Read the interview methodology note
- Finish Kleppmann lectures 1 to 8
- Database selection by requirements
- Partitioning and sharding: cross-partition transactions still unclear
- Three timed mock interviews
- Write up one real design decision from work
Understanding Your Starting Point
Before choosing a learning path, assess your current position:
Background Assessment
Academic Foundation (Strong CS Theory)
- Formal computer science education
- Coursework in algorithms, data structures, databases
- Exposure to distributed systems concepts
- Comfort with academic papers and theoretical frameworks
Practical Experience (Industry Background)
- 2+ years of software development
- Experience with production systems
- Understanding of basic web architectures
- Familiarity with databases, APIs, and deployment
Career Transition (Bootcamp/Self-Taught)
- Strong programming fundamentals
- Limited exposure to system design concepts
- Focused on practical skills over theory
- Timeline pressure for immediate competency
Beginner (New to System Design)
- Basic programming knowledge
- Minimal exposure to distributed systems
- Need comprehensive foundation building
- Long-term learning commitment available
Goal Assessment
Interview Preparation (8-16 weeks)
- Immediate hiring process pressure
- Focus on communication and methodology
- Pattern recognition over deep theory
- Performance under time constraints
Senior Role Transition (12-20 weeks)
- Career advancement within organization
- Leadership and mentorship responsibilities
- Business impact and cost considerations
- Cross-functional collaboration needs
Deep Expertise Building (6+ months)
- Passion for distributed systems theory
- Research and innovation interests
- Thought leadership aspirations
- Long-term skill investment
Practical Implementation (Variable timeline)
- Current work project requirements
- Specific technology decisions needed
- Immediate problem-solving focus
- Learning through application
Learning Path 1: Interview-First Track
“From Zero to Interview-Ready in 12 Weeks”
Target Audience: Career changers, bootcamp graduates, engineers with interview pressure Timeline: 8-12 weeks intensive Primary Goal: Pass system design interviews at major tech companies
Phase 1: Foundation and Methodology (Weeks 1-3)
Build systematic problem-solving approach
Week 1: Process Mastery
- Primary: System Design/system_design_interview_methodology - Master the 6-phase framework
- Practice: Simple problems with methodology focus (URL shortener, chat application)
- Outcome: Consistent time management and structured approach
Week 2-3: Core Concepts
- Primary: System Design/general_lessons_sd_interviews - Understand scalability and consistency fundamentals
- Secondary: System Design/trade-offs - Learn to articulate trade-offs clearly
- Practice: Apply concepts to methodology practice
- Outcome: Technical vocabulary and reasoning ability
Phase 2: Pattern Recognition (Weeks 4-7)
Develop architectural intuition through repetition
Week 4-5: Database Decision Making
- Primary: System Design/database/choosing_database_by_requirements - Systematic database selection
- Secondary: System Design/data_partitioning_and_sharding - Introduction to scaling techniques
- Practice: Database selection for different system types
- Outcome: Confident data architecture decisions
Week 6-7: System Architecture Patterns
- Practice: Work through system design katas systematically
- Focus: Social media, e-commerce, file storage, chat systems
- Method: 45-minute time-boxed sessions
- Outcome: Recognition of common architectural patterns
Phase 3: Advanced Scaling (Weeks 8-10)
Handle complex scalability challenges
Week 8-9: Partitioning Deep-Dive
- Primary: System Design/data_partitioning_and_sharding - Complete partitioning mastery
- Practice: Complex multi-service systems requiring sharding
- Focus: Cross-partition consistency and operation complexity
- Outcome: Ability to design systems handling massive scale
Week 10: Integration and Trade-offs
- Primary: System Design/trade-offs - Advanced trade-off analysis
- Practice: Defend architectural decisions under questioning
- Focus: Alternative solutions and their implications
- Outcome: Sophisticated architectural reasoning
Phase 4: Interview Simulation (Weeks 11-12)
Polish communication and performance under pressure
Week 11-12: Mock Interview Practice
- Method: Timed mock interviews with feedback
- Focus: Communication clarity, whiteboard skills, handling pushback
- Scope: Cover broad range of system types and complexity levels
- Outcome: Interview-day confidence and performance
Success Metrics:
- Consistently complete 45-minute design sessions
- Articulate clear trade-offs for all major decisions
- Handle follow-up questions with confidence
- Design systems supporting 100M+ users
Learning Path 2: Academic Foundation Track
“Deep Theory to Practical Mastery in 16 Weeks”
Target Audience: CS graduates, engineers who prefer theory-first learning Timeline: 14-18 weeks comprehensive Primary Goal: Build deep expertise applicable to senior engineering roles
Phase 1: Theoretical Foundations (Weeks 1-6)
Establish rigorous understanding of distributed systems principles
Week 1-3: Core Distributed Systems Theory
- Primary: Martin Kleppmann Distributed Systems Lectures 1-8
- Focus: Networking, fault tolerance, time and causality
- Method: Lecture + notes + concept verification through simple exercises
- Outcome: Solid foundation in distributed systems fundamentals
Week 4-6: Consensus and Consistency
- Primary: Martin Kleppmann Lectures 9-16
- Focus: Consensus algorithms, replication, consistency models
- Secondary: System Design/Martin Kleppmann Distributed Systems Lecture/6_1_consensus
- Outcome: Deep understanding of correctness in distributed systems
Phase 2: Integration with Practice (Weeks 7-10)
Connect theory to real-world application
Week 7-8: Scaling Principles
- Primary: System Design/data_partitioning_and_sharding - Theoretical understanding applied
- Secondary: System Design/general_lessons_sd_interviews - Practical implications
- Practice: Design partitioning strategies for different data access patterns
- Outcome: Theory-backed scaling decisions
Week 9-10: System Design Methodology
- Primary: System Design/system_design_interview_methodology - Structured application
- Practice: Complex system designs with theoretical justification
- Focus: Explaining “why” as much as “what”
- Outcome: Ability to teach and mentor others
Phase 3: Advanced Topics (Weeks 11-14)
Explore cutting-edge concepts and emerging patterns
Week 11-12: Database Theory and Practice
- Primary: System Design/database/choosing_database_by_requirements
- Secondary: Advanced Martin Kleppmann lectures on storage
- Practice: Design storage solutions for complex requirements
- Outcome: Database architecture expertise
Week 13-14: Case Study Deep-Dives
- Method: Choose 2-3 complex real-world systems (Google Spanner, Amazon DynamoDB)
- Focus: Academic paper analysis + practical implementation considerations
- Practice: Reverse-engineer existing systems
- Outcome: Understanding of production-scale complexity
Phase 4: Teaching and Leadership (Weeks 15-16)
Develop expertise communication and mentorship skills
Week 15-16: Knowledge Synthesis
- Method: Create technical presentations on specialized topics
- Practice: Mentor others through system design problems
- Focus: Simplifying complex concepts for different audiences
- Outcome: Technical leadership readiness
Success Metrics:
- Explain any distributed systems concept from first principles
- Design novel solutions to new problem classes
- Mentor others effectively through complex problems
- Contribute to technical decision-making at organizational level
Learning Path 3: Practical Implementation Track
“Working Engineer to System Architect in 14 Weeks”
Target Audience: Experienced developers transitioning to architecture roles Timeline: 12-16 weeks with work integration Primary Goal: Apply system design knowledge to current projects and advance career
Phase 1: Systematic Knowledge Building (Weeks 1-4)
Formalize intuitive knowledge and fill gaps
Week 1-2: Framework and Assessment
- Primary: System Design/system_design_interview_methodology - Structure existing knowledge
- Assessment: System Design/system-design-skill-ladder - Identify current level and gaps
- Practice: Apply methodology to current work challenges
- Outcome: Systematic approach to architectural decisions
Week 3-4: Trade-offs and Decision Making
- Primary: System Design/trade-offs - Formalize cost-benefit analysis
- Secondary: System Design/general_lessons_sd_interviews - Connect to scaling patterns
- Practice: Document architectural decisions in current projects
- Outcome: Improved decision justification and documentation
Phase 2: Advanced Technical Skills (Weeks 5-8)
Build expertise in areas likely missed in day-to-day work
Week 5-6: Data Architecture Mastery
- Primary: System Design/data_partitioning_and_sharding - Learn systematic partitioning
- Secondary: System Design/database/choosing_database_by_requirements
- Practice: Analyze current data architecture for optimization opportunities
- Outcome: Advanced database and scaling expertise
Week 7-8: Distributed Systems Theory
- Primary: Selected Martin Kleppmann lectures (consensus, replication, consistency)
- Focus: Areas not encountered in current work
- Practice: Identify consistency issues in current systems
- Outcome: Theoretical backing for practical experience
Phase 3: System Design Practice (Weeks 9-12)
Apply knowledge to increasingly complex problems
Week 9-10: Design Exercise Practice
- Method: Work through system design katas with time constraints
- Focus: Systems outside current domain expertise
- Practice: Document design decisions and trade-offs
- Outcome: Broadened architectural perspective
Week 11-12: Cross-Functional Integration
- Practice: Design systems considering team structure, budget, timeline
- Focus: Business requirements translation to technical architecture
- Method: Mock stakeholder presentations and requirements gathering
- Outcome: Business-aligned technical leadership
Phase 4: Leadership Transition (Weeks 13-14)
Develop architecture communication and influence skills
Week 13-14: Architecture Communication
- Practice: Present architectural proposals to technical and non-technical audiences
- Focus: Risk communication, timeline estimation, cost modeling
- Method: Internal architecture reviews and technical design documents
- Outcome: Ready for staff/principal engineer responsibilities
Success Metrics:
- Lead architecture decisions on current team
- Successfully mentor junior engineers through design decisions
- Communicate technical complexity to business stakeholders
- Identify and drive architectural improvements in existing systems
Learning Path 4: FAANG Interview Intensive
“Elite Interview Performance in 10 Weeks”
Target Audience: Experienced engineers targeting top-tier companies Timeline: 8-12 weeks focused preparation Primary Goal: Excel in system design interviews at FAANG-level companies
Phase 1: Rapid Foundation Building (Weeks 1-2)
Establish minimum viable competency quickly
Week 1: Methodology Mastery
- Primary: System Design/system_design_interview_methodology - Perfect the 6-phase approach
- Practice: 3-4 timed practice sessions daily
- Focus: Time management and systematic coverage
- Outcome: Consistent 45-minute performance
Week 2: Core Pattern Recognition
- Primary: System Design/general_lessons_sd_interviews
- Practice: Common system types (social media, chat, file storage)
- Focus: Standard architectural patterns and scaling techniques
- Outcome: Rapid pattern recognition and application
Phase 2: Advanced Technical Depth (Weeks 3-5)
Build sophisticated technical reasoning ability
Week 3: Database and Storage Excellence
- Primary: System Design/database/choosing_database_by_requirements
- Secondary: System Design/data_partitioning_and_sharding
- Practice: Database selection defense under questioning
- Outcome: Confident data architecture decisions
Week 4-5: Complex System Integration
- Practice: Multi-service systems with complex requirements
- Focus: Cross-cutting concerns (security, monitoring, cost)
- Method: Handle increasingly difficult follow-up questions
- Outcome: Sophisticated system reasoning
Phase 3: Communication Excellence (Weeks 6-8)
Perfect interview communication and performance
Week 6-7: Explanation and Teaching
- Practice: Explain designs to different audience levels
- Focus: Clarity, logical flow, handling interruptions
- Method: Record practice sessions for self-review
- Outcome: Clear, confident communication
Week 8: Stress Testing and Edge Cases
- Practice: Handle hostile or challenging interview scenarios
- Focus: Performance under pressure, graceful failure handling
- Method: Mock interviews with aggressive questioning
- Outcome: Unflappable interview performance
Phase 4: Company-Specific Preparation (Weeks 9-10)
Tailor preparation to specific company cultures and expectations
Week 9-10: Company Research and Practice
- Research: Study target companies’ architectural principles and public systems
- Practice: Company-specific system design problems and evaluation criteria
- Focus: Cultural fit and company-specific terminology
- Outcome: Optimized performance for target companies
Success Metrics:
- Complete any system design problem within 45 minutes
- Handle aggressive follow-up questions confidently
- Explain complex trade-offs to any audience level
- Demonstrate knowledge of real-world implementation details
Specialized Focus Areas
Security-First System Design
For engineers focused on secure system architecture
Core Path Integration:
- Complete Interview-First or Academic Foundation track
- Add security considerations to every design decision
- Study threat modeling and security architecture patterns
- Practice secure design under compliance requirements
Additional Resources:
- OWASP architecture guidelines
- Security architecture case studies
- Compliance framework integration (SOC2, HIPAA, GDPR)
Cost-Optimized Architecture
For engineers in cost-sensitive environments or fintech
Core Path Integration:
- Complete Practical Implementation track
- Add cost modeling to every architectural decision
- Study cloud economics and resource optimization
- Practice designing within strict budget constraints
Additional Resources:
- Cloud cost optimization techniques
- Resource utilization monitoring and optimization
- Financial modeling for technical decisions
Observability and Operations
For engineers transitioning to SRE or platform roles
Core Path Integration:
- Complete Academic Foundation track
- Add monitoring and alerting to every system design
- Study incident response and chaos engineering
- Practice designing systems for operational excellence
Additional Resources:
- SRE principles and practices
- Monitoring and alerting strategy
- Incident response and post-mortem culture
Learning Support and Community
Study Groups and Peer Learning
Maximize learning through collaboration
Mock Interview Partnerships:
- Partner with someone at similar level for regular practice
- Alternate interviewer/interviewee roles
- Focus on feedback and improvement areas
- Schedule weekly sessions for consistency
Study Group Formation:
- 3-4 people following similar learning paths
- Weekly discussion of challenging concepts
- Collaborative problem-solving on complex systems
- Shared resource discovery and knowledge synthesis
Progress Tracking and Adjustment
Weekly Self-Assessment:
- Rate comfort level with key concepts (1-10 scale)
- Identify areas of confusion or difficulty
- Adjust timeline and focus based on progress
- Seek additional resources for challenging areas
Milestone Checkpoints:
- Complete practice problems without references
- Explain concepts to others clearly
- Handle increasingly complex scenarios
- Demonstrate consistent performance under time pressure
Transition Between Paths
Path Switching Guidelines:
- Interview → Academic: Add theoretical depth after immediate needs met
- Academic → Interview: Focus on time management and communication
- Practical → Interview: Emphasize methodology and pattern recognition
- Any → Specialized: Build specialized expertise on solid foundation
Conclusion: Choosing Your Journey
System design mastery is not a destination but a continuous journey of learning and application. The paths outlined here provide structure and progression, but your unique background, goals, and constraints should guide your specific choices.
Key Success Factors:
- Consistent Practice: Regular, focused study sessions with practical application
- Active Learning: Explain concepts to others, teach what you learn
- Real Application: Connect learning to current work and projects
- Community Engagement: Learn with others and seek feedback
- Iteration and Adaptation: Adjust your path based on progress and changing goals
Remember: The goal isn’t to memorize solutions but to develop systematic thinking, clear communication, and deep understanding of trade-offs that enable you to design systems that solve real problems effectively.
Choose the path that best fits your starting point and goals, but don’t hesitate to adapt as you progress. The journey to system design mastery is highly personal—these paths provide the roadmap, but you’ll navigate your own unique route to expertise.
Ready to start? Choose your path, commit to the process, and begin your transformation from someone who uses systems to someone who architects them.
Quartz 5 aims at full Obsidian compatibility. This note uses each construct once, with real content, so the rendering can be checked page by page.
Callouts
Plain callout
The default type. Useful for asides that should not interrupt the flow of the argument.
Collapsed by default (click the title)
A dash after the type collapses the callout. Good for long reference material: the reader opts in.
Nested callouts
Callouts can contain callouts.
Inner example
A quorum of 2 out of 3 nodes tolerates exactly one failure. A quorum of 3 out of 5 tolerates two.
Open question
Does a stacked-pages layout make long notes harder to read on a laptop screen? Decide after a week of use.
Highlights and hidden comments
The claim to remember from the CAP discussion is that partition tolerance is not optional; the real decision is what to sacrifice while the network is broken.
Task lists with custom markers
- Migrate the site to Quartz 5
- Review the stacked-pages experience on desktop
- Decide whether the Tokyo Night theme should stay
- Install the Giscus app on the repository so comments can be posted
- Hard line breaks plugin: rejected, the notes are hard-wrapped
- Revisit this list in a month
Block reference
The definition below is not copied; it is embedded from the partition tolerance article through its block id, so it stays in sync:
Transclude of system-design/partition_tolerance_in_distributed_systems#^cap-p-definition
Footnotes
Mermaid renders at build time with an expand button1, and the video below is embedded with plain image syntax2.
Mermaid
flowchart LR A[Markdown note] --> B[Transformers] B --> C[Filters] C --> D[Emitters] D --> E[(public/)] B -. plugins .-> B
Video embed
Martin Kleppmann’s first distributed systems lecture, which the System Design notes on this site follow:
Wikilinks with aliases and headings
- Alias: the CP systems note
- Heading link: System Design/cp_systems_design > What a quorum write looks like
- Tag link: system-design
Footnotes
| Note | Tags | Updated |
|---|---|---|
| Area concept11 | ||
| Dynamic Programming | system-design, interview-prep, mock-interview, design-exercise, reservations, geospatial, intermediate, tutorial | 14m ago |
| Availability in CAP Theorem: Always On, Always Responding | distributed-systems, cap-theorem, availability, databases, system-design | 650d ago |
| Consistency in CAP Theorem: Ensuring Everyone Sees the Same Truth | distributed-systems, cap-theorem, consistency, databases, system-design | 650d ago |
| Data Partitioning and Sharding - The Foundation of Scalable Systems | system-design, distributed-systems, sharding, partitioning, scalability, databases, data-architecture, horizontal-scaling, consistency, performance | 14m ago |
| General Lessons for System Design Interviews | system-design, distributed-systems, concurrency, interviews, scalability, load-balancing, message-queues, databases, cap-theorem, durability, consistency, asynchronous-processing | 14m ago |
| Partition Tolerance in CAP Theorem: The Inevitable Necessity | distributed-systems, cap-theorem, partition-tolerance, network-partitions, system-design | 650d ago |
| System Design Skill Ladder From 30-Kyu to Dan Mastery | system-design, career, learning-path, skill-development, progression, beginner, intermediate, advanced, guide | 14m ago |
| System Design Interview Methodology - A Step-by-Step Framework | system-design, interviews, methodology, framework, communication, process, time-management, scalability, distributed-systems | 14m ago |
| System Design Learning Paths - Your Journey to Distributed Systems Mastery | system-design, learning-paths, career-development, distributed-systems, interviews, methodology, skill-progression, study-guide | 14m ago |
| System Design Study Guide | system-design, study-guide, learning-path, interview-prep, ai-assistant, practice, beginner, intermediate, advanced, scalability, distributed-systems, ai-tools, mermaid, database-design, api-design | 14m ago |
| System Design trade-offs | system-design, trade-offs, cap-theorem, performance, consistency, scalability, intermediate, guide | just now |
| Area database4 | ||
| Choosing the Right Database | system-design, databases, data-modeling, performance, comparison, intermediate, guide | just now |
| Database Indexing | system-design, databases, performance, indexing, data-structures, advanced, guide | just now |
| Why Use DynamoDB for a Small Project? | system-design, databases, nosql, sql, comparison, aws, intermediate, guide | just now |
| Integrating and Normalizing Data from Multiple Sources | system-design, databases, data-modeling, integration, normalization, architecture, intermediate, guide | 14m ago |
| Area kata4 | ||
| 4 additional system design katas | system-design, kata, exercise, practice, problem-solving, intermediate, tutorial | 14m ago |
| Blurry to Sharp Technique | system-design, kata, exercise, practice, methodology, collaborative-systems, intermediate, tutorial | 14m ago |
| System Design exercises - idealized scenarios | system-design, kata, exercise, practice, problem-solving, theoretical, intermediate, tutorial | 14m ago |
| System Design, Building a Real-Time Collaborative Editor | system-design, kata, exercise, practice, real-time, collaborative-systems, intermediate, tutorial | 14m ago |
| Area Kleppmann lecture23 | ||
| 1.1 Distributed Systems Lecture - Intro | system-design, distributed-systems, martin-kleppmann, intermediate, reference | 14m ago |
| 1.2 Computer Networking | system-design, distributed-systems, martin-kleppmann, intermediate, reference, networking, protocols | 14m ago |
| 1.3 Remote Procedure Call | system-design, distributed-systems, martin-kleppmann, intermediate, reference, networking, protocols | 14m ago |
| 2.1 The Two generals problem | system-design, distributed-systems, martin-kleppmann, intermediate, reference, fault-tolerance, reliability | 14m ago |
| 2.2 The Byzantine generals problem | system-design, distributed-systems, martin-kleppmann, intermediate, reference, fault-tolerance, reliability | 14m ago |
| 2.3 System Models | system-design, distributed-systems, martin-kleppmann, intermediate, reference, fault-tolerance, reliability | 14m ago |
| 2.4 Fault tolerance | system-design, distributed-systems, martin-kleppmann, intermediate, reference, fault-tolerance, reliability | 14m ago |
| 3.1 Physical Time | system-design, distributed-systems, martin-kleppmann, intermediate, reference, time, synchronization | 14m ago |
| 3.2 Clock Synchronisation | system-design, distributed-systems, martin-kleppmann, intermediate, reference, time, synchronization | 14m ago |
| 3.3 Causality and happens-before | system-design, distributed-systems, martin-kleppmann, intermediate, reference, time, synchronization | 14m ago |
| 4.1 Logical time | system-design, distributed-systems, martin-kleppmann, intermediate, reference, time, synchronization | 14m ago |
| 4.2 Broadcast Ordering | system-design, distributed-systems, martin-kleppmann, intermediate, reference, broadcast, ordering | 14m ago |
| 4.3 Broadcast Algorithms | system-design, distributed-systems, martin-kleppmann, intermediate, reference, broadcast, ordering, algorithms | 14m ago |
| 5.1 Replication | system-design, distributed-systems, martin-kleppmann, intermediate, reference, replication, consistency | 14m ago |
| 5.2 Quorums | system-design, distributed-systems, martin-kleppmann, intermediate, reference, replication, consistency | 14m ago |
| 5.3 State machine replication | system-design, distributed-systems, martin-kleppmann, intermediate, reference, replication, consistency | 14m ago |
| 6.1 Consensus | system-design, distributed-systems, martin-kleppmann, intermediate, reference, consensus, algorithms | 14m ago |
| 6.2 Raft | system-design, distributed-systems, martin-kleppmann, intermediate, reference, consensus, algorithms | 14m ago |
| 7.1 Two-phase commit | system-design, distributed-systems, martin-kleppmann, intermediate, reference, transactions, commit-protocols | 14m ago |
| 7.2 Linearizability | system-design, distributed-systems, martin-kleppmann, intermediate, reference, consistency, linearizability | 14m ago |
| 7.3 Eventual consistency | system-design, distributed-systems, martin-kleppmann, intermediate, reference, consistency | 14m ago |
| 8.1 Collaboration software | system-design, distributed-systems, martin-kleppmann, intermediate, reference, crdt, collaboration | 14m ago |
| 8.2 Google's spanner | system-design, distributed-systems, martin-kleppmann, intermediate, reference, google-spanner, databases | 14m ago |
| Area lab2 | ||
| CAP Theorem Reading List | test, quartz-v5, system-design, distributed-systems | 14m ago |
| Obsidian Syntax Showcase | test, quartz-v5, obsidian, system-design | 14m ago |
| Area mock interview3 | ||
| App store design | system-design, interview-prep, mock-interview, design-exercise, scalability, file-handling, intermediate, tutorial | just now |
| Geohashing | system-design, interview-prep, mock-interview, design-exercise, reservations, geospatial, intermediate, tutorial | 14m ago |
| Designing a Parking Garage Reservation System | system-design, interview-prep, mock-interview, design-exercise, reservations, geospatial, intermediate, tutorial | just now |
Pan by dragging, zoom with the wheel, reset with the control on the right. On phones the sidebar overlays the canvas.