Skip to content
Backend & systems

API Idempotency Keys: Safe Retries Under Concurrency

A client times out after creating a booking. It cannot tell whether the server committed the booking or never received the request. Retrying with a new operation identity can create another booking. An idempotency key lets the server recognize a retry, but only a defined server contract turns that key into protection.

By Published

4 min readEditorial analysisUpdated
  • API design
  • Idempotency
  • Retries
  • Concurrency
What to remember

Keep one key per intended operation, validate its payload, and claim it atomically. Describe concurrency, retention, and failure behavior explicitly; a header alone cannot coordinate side effects.

Identify intent, not merely identical input

Two identical bookings can be intentional; two attempts to create one booking are a retry. AWS's design discussion explains why caller-supplied operation identifiers are useful for this distinction. Generate the key before the first attempt and preserve it across retries. A changed booking is a new operation with a new key.

In our modeled booking API, scope the key to the authenticated customer, method, and route. Store a fingerprint of validated, normalized business input. Do not derive customer identity from an untrusted request body.

Define the duplicate-request contract

These are choices for our example API, not mandatory status codes for every provider. A repeated key with different input must not silently return an unrelated booking. Define a bounded waiting policy for an operation already running and publish how long completed results remain replayable.

Our booking API's modeled behavior
RequestServer decision
New scoped keyClaim and execute
Completed key, same inputReplay stored result
Same key, changed inputReject key reuse
Matching request still runningWait briefly or report conflict

Claim before executing the booking

A lookup followed by an insert is racy: two workers can both observe no record. For a database-contained operation, use a unique constraint on the scoped key and attempt the insert atomically. PostgreSQL's ON CONFLICT supports this claim pattern. Only the worker that obtains the claim performs the booking mutation.

The following is design pseudocode, not executable SQL. Keep the claim, booking mutation, and saved result in one transaction. At PostgreSQL's Read Committed isolation, a competing claim can wait for another transaction; read the completed record in a subsequent statement after a conflict. Handle lock timeouts explicitly.

BEGIN TRANSACTION
try INSERT claim with UNIQUE(customer, method, route, key)
if claim inserted:
  create booking
  save input fingerprint and response
  COMMIT
else:
  read existing claim in a subsequent statement
  reject changed input, or replay its completed response
  COMMIT

The transaction boundary sets the guarantee

If this transaction rolls back, neither the booking nor its successful result remains. If it commits and the connection drops before the response arrives, the matching retry reads the saved result. That resolves the booking example's ambiguous timeout.

An external charge or email is outside that database transaction. A crash after the external action but before recording success leaves uncertainty. Use the external provider's documented idempotency mechanism and reconciliation where needed. Publishing a single key does not establish an end-to-end exactly-once guarantee.

Read the provider contract before retrying

Stripe documents replaying the first saved status and body, including 500 responses. It also documents parameter comparison, possible key removal after at least 24 hours, and cases where validation or concurrent conflicts do not create a saved result. Those are Stripe's rules, not properties of every idempotency implementation.

Keep the original key during an uncertain retry, follow the provider's retry guidance, and bound attempts with backoff. After retention expires, do not assume an old key still prevents a new operation. Reconcile unresolved outcomes instead of generating new keys until something succeeds.

Test the failure timeline

Use a disposable service and database for this exercise. Inspect booking records as well as responses. A test that only checks HTTP success can miss duplicate writes or a result recorded before the booking commits.

  • Send simultaneous matching requests and verify one booking.
  • Repeat the key with changed input and verify rejection.
  • Drop the response after commit, then retry the original key.
  • Test rollback, timeout, and expired-key behavior separately.

Quick answers

Frequently asked questions

Is a new key needed for each retry?

No. Retrying the same intended operation uses the original key. A genuinely new operation gets a new key, even when its business input happens to match.

Can a cache lookup prevent concurrent duplicates?

A separate lookup cannot atomically claim work. Use a claim mechanism with enforced uniqueness and define what competing workers do while the first operation is running.

Does an idempotency key guarantee delivery?

No. It supports duplicate handling within a stated contract. Availability, retry limits, retention, and external effects still determine whether the operation completes and how uncertainty is resolved.

Source notes

References and review policy

Information checked on October 4, 2026. Section links identify sources for factual claims and technical explanations. Interpretations, practice scenarios and preparation recommendations are RecallDeck’s editorial work.

From reading to recall

Practice the full interview loop.

RecallDeck schedules the concepts you miss and keeps coding, design, and behavioral fundamentals available when the interviewer changes direction.

Start studying

Keep going