Marcio Cunha

Processing Long-Running Distributed Transactions with Orchestrated Sagas and Event Sourcing

Learn how to coordinate complex operations across microservices while maintaining eventual consistency without locking databases. A practical guide on the Saga pattern combined with immutable event logs.

Marcio Cunha•4 min
Also available in:EspañolPortuguês
Summary
  • Modern distributed systems require abandoning classic database transactions in favor of event-driven eventual consistency.
  • Choosing between choreography and orchestration dictates whether compensation logic is scattered across services or centralized in a dedicated component.
  • Immutable state logging guarantees complete auditing and safe fault replay without historical data loss.
  • Long-running transactions must account for network partitions and temporary third-party outages from the design phase.
  • Proper compensation workflows ensure systems return to a consistent state even after partial failures in advanced steps.

The Challenge of Distributed Transactions in Microservices

When splitting a giant monolithic system into smaller pieces called microservices, we gain the freedom to scale each function independently. In a monolith, ensuring a purchase happens entirely from charging the credit card to reserving stock is simple because everything uses the same database with an atomic transaction instruction. In practice, this means either everything works together or nothing changes. In the distributed world, each service owns its isolated database, making it impossible to use this native facility to unite operations that cross distinct server boundaries.

To solve this problem without relying on slow locks that freeze the entire application, engineers adopt the Saga pattern. A Saga is a sequence of local transactions where each step updates data in a specific service and publishes an event announcing task completion. If something goes wrong halfway through, the Saga executes compensating transactions, which practically function as a logical undo of what was already done, such as refunding a payment or returning an item to stock. This model trades immediate consistency for eventual consistency, accepting brief moments of misalignment until all steps finish successfully.

Choosing Between Choreography and Orchestration in Sagas

There are two main ways to implement the Saga pattern: choreography and orchestration. In choreography, there is no central boss controlling the flow; each microservice listens to events generated by others and autonomously decides the next step. In practice, it is like a ballroom dance where each participant reacts to the other's movement without needing a choreographer shouting orders. Although it works well for short flows, choreography can quickly turn into a tangled web that is hard to trace when business processes involve dozens of steps and complex conditional rules.

This is where orchestration comes into play. In the orchestrated model, there is a dedicated component called an orchestrator, whose sole responsibility is to dictate the rhythm and order of calls. In practice, the orchestrator receives the initial request, sends a command to the first service, awaits the response, notes the result, and dispatches the next command. If any step fails, the orchestrator consults its internal log and triggers compensation commands in the exact reverse order. This centralization brings total process visibility, easing error debugging and the addition of new business rules without altering domain services.

Integrating Immutable Event Logging with the Orchestrator

For the orchestrator to make safe decisions, it must know precisely what state the process is in at every millisecond. This is where Event Sourcing comes in, which consists of storing all state changes of the system as a chronological sequence of immutable events rather than just keeping the current table snapshot in the database. In practice, the orchestrator's database does not record that an order is in the 'paid' stage, but rather the exact list of everything that happened: 'order created', 'payment approved', 'stock reserved'. This log acts like an application bank statement, allowing the reconstruction of any Saga's state from scratch simply by replaying history.

Combining the orchestrator with Event Sourcing brings impressive resilience to high-volume architectures. If the orchestrator server experiences a sudden power outage during a complex flow, it does not lose track upon reboot. It simply reads the last recorded events and continues from where it left off, without duplicating charges or leaving orphaned orders. Furthermore, this approach offers a complete and transparent audit trail, essential for meeting rigorous regulatory requirements and facilitating performance analysis and operational bottlenecks in real time.

Handling Failures and Long-Running Transactions

Real business processes frequently involve steps that take hours or even days to complete, such as manual credit verification or physical delivery of goods. In long-running transactions, keeping connections open is unfeasible and dangerous, forcing the architecture to operate in a fully asynchronous manner. In practice, this means the orchestrator dispatches a task, releases machine resources, and enters a passive waiting state until the response event arrives through a message bus like Apache Kafka or RabbitMQ.

The major challenge of these prolonged operations is dealing with unpredictable external failures, such as the temporary downtime of a partner carrier's system. The orchestrator must implement robust strategies for smart retries with increasing intervals, alongside defining maximum tolerance thresholds for each step. If the partner fails to respond within the expected window, the orchestrator assumes failure and triggers the compensation workflow to undo previous reservations. This discipline turns chaotic flows into predictable processes, ensuring software operates reliably even when the surrounding world fails.

Final Considerations on Event-Driven Architectures

Adopting the orchestrated Saga pattern alongside immutable event logging requires a profound shift in mindset for software developers and architects. Abandoning the comfort of traditional database transactions is daunting at first, but the rewards in terms of scalability, resilience, and team independence are unmatched. In practice, systems built on this foundation can absorb massive traffic peaks without crashing core services, gracefully containing failures through well-planned automatic compensations and ensuring long-term data integrity.