Hello! Let's dive into our next topic.
Introduction
In our last lesson, we explored Command Query Responsibility Segregation (CQRS), a pattern that separates the models for writing data (Commands) from those for reading data (Queries). We saw how this allows for independent optimization, leading to a robust domain model on the write side and a highly efficient, denormalized model on the read side.
Today, we'll examine a powerful pattern that often goes hand-in-hand with CQRS and fundamentally changes how we think about data persistence. The learning outcome for this lesson is to: Explain Event Sourcing concepts: event stream, aggregates, projections, and snapshots.
Instead of storing only the current state of your data (e.g., a user's current address in a database row), Event Sourcing involves storing the full, immutable history of state changes as a sequence of events (e.g., UserRegistered, AddressChanged, SubscriptionUpgraded). The current state becomes a derivative of this history, not the primary source of truth.
This approach has profound implications for building complex, auditable, and scalable systems, which is highly relevant to your work in finance and high-load platforms. The concept of an immutable ledger is central to accounting and trading, and Event Sourcing brings that same principle into software architecture.
1. The "Why" of Event Sourcing: From State to History
Before we define the technical components, it's essential to understand the motivation behind Event Sourcing. Why would we choose this more complex method over a simple UPDATE statement? Often, the need arises from business requirements that are difficult to satisfy with a traditional CRUD (Create, Read, Update, Delete) model.
The following video tells a short story about a developer who discovers the need for Event Sourcing organically when faced with questions about past system behavior—questions his state-based persistence model couldn't answer.
Things I wish I knew before I started with event sourcing By Michał Ostruszka
This video, 'Things I wish I knew before I started with event sourcing' by Michał Ostruszka, provides an excellent narrative introduction. It illustrates how business needs for auditing and historical data naturally lead to the idea of logging events as the source of truth.
Please watch the introductory story (02:19 - 10:00). As you watch, consider how the developer's problem—needing to answer questions about past events that were never recorded—is a common challenge in analytics, debugging, and compliance.
As the video highlights, the core idea is a shift in perspective: from viewing the system as a collection of current states to viewing it as a sequence of historical facts. This not only provides a rich audit log but also unlocks powerful capabilities for analytics and system design.
2. The Core Concepts of Event Sourcing
Now, let's formally define the key components that make up an Event Sourcing system. We will use a combination of text and video resources to build a clear picture of each concept.
First, a concise textual overview will help frame the components.
Understanding Event Sourcing with Marten
This guide from MartenDB, 'Understanding Event Sourcing,' provides clear, succinct definitions of the core concepts we are about to discuss.
Please read the 'Core Concepts of Event Sourcing' section. It provides a quick reference for the four key terms: Events, Aggregates/Event Streams, Projections, and Snapshots. We will explore each of these in more detail.
With those definitions in mind, let's explore each concept in depth.
a) Events and the Event Stream
- Event: An immutable record of a business-significant occurrence. Events are named in the past tense (e.g.,
OrderPlaced,PaymentProcessed). They are the atomic unit of change and the single source of truth. - Event Stream: An ordered sequence of events for a single business entity. For example, all events related to a specific customer account (
customer-123) would form a single stream.
The following video segment explains these fundamental terms with a concrete example.
Event Sourcing You are doing it wrong by David Schmitz
In this talk, 'Event Sourcing You are doing it wrong' by David Schmitz, the speaker breaks down the basic terminology of Event Sourcing, connecting it to Domain-Driven Design (DDD).
Watch the section on 'Event Sourcing Terms' (04:31 - 06:40). Pay attention to how he distinguishes between events, streams, and the event store (the persistence mechanism).
b) Aggregates
In Event Sourcing, the Aggregate (a concept from DDD representing a transactional consistency boundary) is not stored directly. Instead, its state is reconstituted on demand.
Here's the process:
- A command is received to change the state of an aggregate (e.g.,
ApproveOrderCommandfororder-456). - The system retrieves the complete event stream for
order-456. - It creates an empty
Orderaggregate and applies each event from the stream to it, one by one (OrderCreated,OrderItemAdded, etc.). This process is often called rehydration. - Once the aggregate is rehydrated to its current state, the command is executed against it. The aggregate's business logic validates the command.
- If the command is valid, the aggregate produces one or more new events (e.g.,
OrderApproved). - These new events are appended to the end of the event stream for
order-456. They are never modified or deleted.
This rehydration process is the heart of the Event Sourcing write model.
Things I wish I knew before I started with event sourcing By Michał Ostruszka
Let's return to the 'Things I wish I knew...' video, which provides a clear animation of this rehydration process.
Watch the segment 'How Event Sourcing Works' (10:00 - 13:09). The animation clearly shows how state is rebuilt from events before a command is processed, and how the new event is appended to the stream.
c) Projections (Read Models)
If the write model is an append-only log of events, how do we query the current state efficiently? We can't rehydrate every aggregate in the system just to find all customers in a specific city.
This is where Projections come in. A projection is a read model that is generated by listening to the event stream and transforming the events into a query-optimized state. This is the direct link back to our previous lesson on CQRS.
- The Event Store is the write database.
- Projections are the read databases.
You can have multiple, independent projections built from the same event stream, each tailored for a different query or UI component. For example:
- An
OrderPlacedevent could update aCustomerDashboardProjection(a SQL table). - The same event could also update an
InventoryProjection(perhaps a Redis cache).
Things I wish I knew before I started with event sourcing By Michał Ostruszka
This segment explains how projections create the read side of a CQRS architecture.
Watch the section on 'Read Models and CQRS' (13:09 - 15:23). Notice how the event stream acts as a source for building multiple, diverse read models, enabling the independent scaling of the read side.
d) Snapshots
A key concern with Event Sourcing, especially in high-load systems, is the performance of rehydrating aggregates. An account with years of transaction history could have millions of events. Replaying them all for every command would be unacceptably slow.
Snapshots are a performance optimization to address this.
- A snapshot is a persisted copy of an aggregate's full state at a specific point in time (i.e., after a certain event number).
- To rehydrate the aggregate, the system loads the most recent snapshot and then replays only the events that have occurred since that snapshot was taken.
This drastically reduces the number of events that need to be processed for long-lived aggregates.
Things I wish I knew before I started with event sourcing By Michał Ostruszka
This final segment on core concepts explains the role and implementation of snapshots.
Watch the 'Performance Optimization with Snapshots' section (44:39 - 47:41). The key takeaway is that snapshots are an optimization, not a core part of the pattern. You should only introduce them when you measure a performance bottleneck.
3. Benefits and Drawbacks
Event Sourcing is a powerful pattern, but it introduces significant complexity. It's crucial to understand the trade-offs before adopting it. Your experience with low-latency trading systems will give you a good intuition for these trade-offs, as they often involve balancing consistency, performance, and auditability.
The 'Pattern: Event sourcing' article on microservices.io provides a balanced summary of the resulting context when you apply this pattern.
Please read the 'Resulting context' section, which lists the benefits and drawbacks. Pay particular attention to the drawbacks, as they highlight the new challenges you must manage: the learning curve, query complexity (necessitating CQRS), and eventual consistency.
Key Benefits:
- Complete Audit Trail: Provides a 100% reliable log of every change, which is invaluable for debugging, compliance, and business intelligence.
- Temporal Queries: You can reconstruct the state of an entity at any point in time.
- Flexibility: New read models (projections) can be created from the event history to answer new business questions without changing the write model.
- Decoupling: Services can subscribe to event streams to react to changes, fostering a loosely coupled, event-driven architecture.
Key Drawbacks:
- Complexity: It's a significant paradigm shift from traditional CRUD.
- Eventual Consistency: Read models are updated asynchronously, so they can be temporarily out of sync with the write model.
- Event Schema Evolution: Once an event is written, it's immutable. Changing the structure of an event ("versioning") requires careful handling strategies.
- Querying: Direct querying of the event store is difficult. CQRS is not just recommended; it's practically a requirement.
Conclusion
Today, we've unpacked the fundamental concepts of Event Sourcing. This pattern offers a robust and auditable way to manage state, especially in complex domains, by treating the history of events as the ultimate source of truth.
Key Takeaways:
- Event Sourcing persists the state of an entity as a sequence of immutable, state-changing events in an Event Stream.
- An Aggregate's current state is not stored directly but is rehydrated by replaying its event stream. It serves as the consistency boundary for validating commands and producing new events.
- Projections are the read models in a CQRS architecture, built by consuming the event stream and creating query-optimized views of the data.
- Snapshots are a performance optimization used to reduce the rehydration time for aggregates with long event streams.
- The primary benefits are a complete audit log and flexibility, while the main challenges are increased complexity and dealing with eventual consistency.
Preview of the Next Lesson
We have now covered the "what" and "why" of Event Sourcing. In our next lesson, "Implement persistence of an aggregate's state as an event sequence in an event store," we will move to the "how." We will look at the practical implementation details of capturing commands, rehydrating aggregates, and appending new events to a persistent event store.