Marcio Cunha

Idempotency and Retries in Webhooks: Handling Duplicate Events

Learn how to build deduplication keys and guarantee at-least-once delivery in webhooks. Find out what to do when partners trigger the same event twice.

Marcio Cunha•5 min
Also available in:EspañolPortuguês
Summary
  • Distributed systems rely on at-least-once delivery because networks fail and automatic retries happen constantly
  • Deduplication keys based on unique identifiers protect databases from processing repeated operations
  • Idempotent processes yield the exact same practical outcome even when executed multiple times with identical data
  • Retention windows in distributed caches prevent concurrent requests from bypassing the initial check
  • Standardized HTTP responses ensure the partner system knows the event was received without needing to resend

The Chaotic Reality of Inter-System Communication

Imagine ordering food through an app and, right as you confirm payment, your internet drops. The system does not know if the order went through or if your money vanished, so it tries sending the instruction again. In software engineering, this network uncertainty is called at-least-once delivery. In practice, this means that to ensure no important message gets lost along the way, servers prefer sending the same notice multiple times rather than risking data loss. Webhooks, which are automated notifications sent from one system to another when something happens, experience this behavior intensely.

When a payment partner notifies you that an invoice was paid, it fires an HTTP request to your server. If your system takes an extra second to respond due to database lag, the partner's server assumes the message never arrived. Immediately, it fires the exact same notice again. Suddenly, your system receives two identical copies of the same payment event within a few milliseconds. If you are not prepared, your code might process the payment twice, generate duplicate ledger entries, or grant double credits to the customer, causing severe operational headaches.

The Vital Concept of Idempotency

To survive this chaos, we must adopt a fundamental engineering concept called idempotency. In mathematics, an operation is idempotent when you can apply it multiple times and the final outcome remains identical to the first execution. Think of an elevator button: if you press the fifth floor button ten times in a row, the elevator will not go to the fiftieth floor; it simply goes to the fifth and stays there. Pressing it a thousand times produces the exact same practical effect as pressing it once.

When building APIs and webhooks, creating an idempotent endpoint means teaching your system to recognize that a current request is identical to one that was already processed successfully in the past. When the second trigger arrives, the system realizes the work is already done, discards the duplicate, and returns a success response as if everything is fine. This eliminates the risk of unwanted side effects, turning a chaotic flood of retries into a predictable, secure, and resilient process.

Implementing Deduplication Keys with Databases

The most straightforward tool to combat duplicate messages is the deduplication key. Every serious webhook provider sends a unique identifier along with the data for that specific event, often called an event ID or idempotency key. When your server receives this package, the first thing it does before touching any business logic is check if that identifier already exists in a control table inside the database.

To make this verification fail-proof, we use uniqueness constraints on database columns. Here is a practical example in Node.js using a relational table:

async function processWebhook(event) {
const { eventId, data } = event;

try {
// Try to register the event in the control table
await db.query(
'INSERT INTO received_events (event_id, status) VALUES ($1, $2)',
[eventId, 'PROCESSING']
);
} catch (error) {
// If there is a unique key violation, the event was already seen
if (error.code === '23505') {
console.log(`Ignored duplicate event: ${eventId}`);
return { status: 200, message: 'Already processed previously' };
}
throw error;
}

// Execute the real business logic safely
await executeBusinessLogic(data);

await db.query(
'UPDATE received_events SET status = $1 WHERE event_id = $2',
['COMPLETED', eventId]
);

return { status: 200, message: 'Successfully processed' };
}

In this code, if two requests arrive at the exact same second, the database rejects the second insert due to the uniqueness constraint on the event_id column. This stops both processes from executing the business logic simultaneously.

Handling Extreme Concurrency Using Distributed Cache

Although database constraints solve most problems, high-scale systems face an extra challenge called race conditions. If two identical triggers arrive nanoseconds apart and the database is still writing the first record, both might pass the initial check before the lock takes effect. To close this gap, modern architectures often use fast in-memory cache systems like Redis.

Redis allows setting temporary locks known as distributed locks or keys with short expiration times. Before querying the relational database, the server tries to write a record with the event ID into the cache using a command that fails if the key already exists. This atomic in-memory operation happens extremely fast, blocking any duplicate attempt before it even touches the main database, protecting expensive infrastructure resources against traffic spikes.

The Role of Retries and HTTP Status Codes

When dealing with webhooks, automatic retries are triggered by the sender when it receives no clear, immediate response. Therefore, how your server responds is crucial to avoiding infinite retry loops. If your code fails due to a temporary error, such as a database connection glitch, you should return an HTTP status code in the 500 range, like 503 Service Unavailable. This tells the partner the issue is on your end and they should try again later.

On the other hand, when a webhook is successfully processed or recognized as a harmless duplicate, your server must return a code in the 200 range, such as 200 OK or 204 No Content. Never return client errors like 400 Bad Request for valid duplicate events unless the message body is actually corrupted. Returning success for a duplicate tells the partner system the message reached its final destination, ending the retry cycle and bringing peace to both servers.

Conclusion

Building webhook integrations requires abandoning the illusion that computer networks are perfectly stable. Understanding that duplicates and delays are part of everyday engineering forces us to design systems focused on resilience, idempotency, and strict concurrency control. By combining smart deduplication keys, robust database constraints, and proper HTTP responses, we turn chaotic events into predictable and secure data flows.

Ultimately, taking care of idempotency is not just a minor programming detail, but an architectural decision that protects the financial and operational integrity of the business. When your systems can absorb duplicate retries without blinking, you gain the peace of mind needed to scale your application without fearing unpleasant surprises in customer records.