Event-driven systems are good at moving work asynchronously, but they introduce an uncomfortable property: you rarely control how many times an event gets delivered. A webhook can be retried. A queue can redeliver a message when a worker crashes. Two workers can process the same message concurrently. The difficult part is not detecting duplicate events. The difficult part is preventing those duplicates from producing duplicate side effects. For example, consider an automation system where a user action eventually triggers an external API call: User Action --> Webhook --> Event Queue --> Worker --> External API If the worker crashes after the external API accepts the request but before our database records the result, we have an ambiguous state. Retrying is necessary for availability, but retrying the external call may create the same side effect again. This is where most of the interesting engineering work begins. Idempotency Has to Be Defined at the Business Level A common implementation is to deduplicate using the provider's event ID: event_id = 123 That is useful, but it is not always enough. The event ID answers: Have I seen this particular event? What we actually need to answer is: Have I already performed this particular business action? Those are different questions. For an automation system, an action might be uniquely identified by: account_id + automation_id + source_event_id + action_type For example: account_123:automation_42:comment_981:send_dm Every retry of the same action must produce the same idempotency key. Generating a new UUID for every attempt defeats the entire purpose: Attempt 1 → key A Attempt 2 → key B The system now sees two operations instead of two attempts at one operation. Don't Use "Check Then Insert" Use application logic to decide what should happen, and database constraints to guarantee what cannot happen. The Hard Failure Case** The real problem appears after the action has been claimed. Consider: Create action Call external API External API succeeds Worker crashes Completion state is never written On recovery: action = processing The system doesn't know whether the external operation happened. This creates the classic distributed-systems gap: Our Database External System commit │ ├──────────────→ API request │ │ │ success │ X worker dies There is no normal database transaction that can atomically commit our database and an unrelated external API. So we need to design around the uncertainty rather than pretending it doesn't exist. Idempotency Keys at the API Boundary If the external API supports idempotency keys, use them. The same logical operation should always carry the same key: action_id = act_12345 Then: First attempt → act_12345 Retry → act_12345 Retry again → act_12345 The external service can safely treat these as attempts for the same operation. When the API doesn't support idempotency, we have to maintain the guarantee ourselves. That usually means storing a durable action record and making the business key unique. The important distinction is that the event is not the unit of idempotency; the side effect is. Queues Don't Solve Duplicate Processing Queues make event-driven systems much easier to scale, but they don't remove this problem. A worker might do: receive message ↓ process ↓ external API succeeds ↓ worker crashes before acknowledgement The queue has no reliable way to know whether the external operation succeeded. It may deliver the message again. Therefore, every queue consumer that performs an external side effect should be safe to retry. Webhook-level deduplication is not enough. Queue-level deduplication is not enough. The operation itself needs to be idempotent. State Machines Are More Useful Than a Boolean A simple: processed = true usually isn't enough. A production action benefits from explicit state: PENDING ↓ PROCESSING ↓ COMPLETED PROCESSING ↓ FAILED ↓ RETRY This lets us distinguish between: work that hasn't started work currently being attempted successfully completed work retryable failures permanently failed work It also gives recovery processes something durable to reason about. For example, a job stuck in PROCESSING for longer than the expected execution window can be investigated or recovered. But recovery must still use the same idempotency mechanism. A stale job should never mean "send the request again without checking." Retries Need Classification Not every failure should trigger another attempt. A timeout or 500 usually represents a transient failure. A 429 generally means we should slow down and retry later. An invalid request or invalid permission is different. Repeating the same request will not fix it. A useful mental model is: Transient failure → retry Rate limit → backoff + retry Permanent failure → stop Unknown failure → investigate safely Exponential backoff is useful here, especially when many workers encounter the same downstream problem. Otherwise a temporary outage can turn into a retry storm. Observability Is Part of Idempotency When duplicate side effects occur, the most important question is often: Which path caused the second execution? That is almost impossible to answer if logs contain only: DM sent We instead want a traceable chain: event_id action_id automation_id account_id attempt worker_id external_request_id status Then we can reconstruct: Event → Action → Attempt → API Request and distinguish: duplicate event from: duplicate action from: retry after unknown API outcome Those are very different failures. Exactly Once Is Usually the Wrong Goal We often hear: "We need exactly-once processing." In practice, guaranteeing exactly-once execution across a webhook provider, queue, database, worker, and external API is extremely difficult. A more useful design is: At-least-once delivery + Idempotent business actions + Durable state + Safe retries The infrastructure may process an event multiple times. The user-visible result should still happen only once. That is the property we actually care about. What We Use in Practice When building event-driven automation, the pattern that works well is: External Event ↓ Validate + persist ↓ Stable event identity ↓ Queue ↓ Create/claim business action ↓ Unique database constraint ↓ Idempotent external request ↓ Update durable state ↓ Metrics + logs The key lesson is simple: Don't try to make the entire system exactly-once. Make every important side effect safe to execute more than once. That shift in thinking changes how you design the database, queues, workers, retries, and API integrations. And once your system starts processing events at scale, that small design decision can be the difference between a retry being a recovery mechanism and a retry becoming a customer-facing bug. We encountered these problems while building event-driven automation at Vyral , where user events can trigger external messaging actions. The same patterns apply to payment processing, notifications, order workflows, and almost any system where an event can create an external side effect.