Retry agent tool actions without duplicating their effects
Use persistent operation identifiers, explicit unknown outcomes and reconciliation to control retries of external actions in LLM integrations.
An agent creates a support ticket. The provider accepts the request, but its response is lost. Your application treats the timeout as failure and sends another create request. The support team now has two tickets for the same intended action.
This hypothetical failure occurs between an effect and its confirmation. Safe recovery requires identifying the original operation and establishing its outcome. A model's subsequent explanation cannot tell the application what the provider actually committed.
Give the business intention a persistent identity
Create an operation identifier when the application accepts a specific action. Keep it across retries and process restarts, including cases where the model proposes the same tool call again. A model-generated call identifier may be new on each execution and is not necessarily a durable business-operation key.
Identical request bodies can represent separate intentions. A user may genuinely want two similar tickets. Hashing the body alone therefore cannot decide whether a request is a duplicate. AWS's idempotent API guidance describes caller-provided identifiers as a way to express that intent.
Scope the identifier to the tenant and authorised operation. Possession of another customer's key must not grant access to their result. Keep the authorisation check separate from duplicate detection.
Persist the command before dispatch
An illustrative operation record might include:
operation_id
tenant_id
action_type
arguments_fingerprint
status
provider_reference
Record the authorising identity and approved proposal version as well as timestamps needed for recovery. An argument fingerprint helps detect a reused key with different contents. It cannot establish whether the caller is allowed to perform the action.
Use a uniqueness constraint and atomic state transitions to prevent concurrent workers from claiming the same operation. A duplicate request can return a stored confirmed result or report that work is still pending. Avoid starting a second independent execution merely because the first worker has not returned.
PostgreSQL documents an atomic insert-or-update outcome for INSERT with ON CONFLICT. That guarantee covers the database operation. The provider call and your database write typically occur in different systems, leaving a failure interval that needs explicit recovery.
Check the provider's actual contract
Where an endpoint supports idempotency keys, send the persistent operation key and reuse it for retries of that action. Read the endpoint's rules for stored responses, changed parameters and retention. These details vary by provider.
A network failure is not a reason to generate a fresh key. An old key whose retention window has expired also needs care: the provider may treat it as a new request. Establish whether the original effect exists before deciding to dispatch again.
If the endpoint lacks idempotency support, look for an external reference that allows status reconciliation. When no reliable lookup exists, preserve the unknown outcome and require human investigation. Blindly repeating a write can turn a temporary connection problem into a permanent duplicate.
Represent uncertainty in the runtime
A validation rejection before execution differs from a lost response after execution. Preserve that distinction in state and in the user-facing message.
Track commands that are ready to send, executing and confirmed. Add an unknown-outcome state with a defined reconciliation path. State who may resolve it and which evidence from the provider establishes completion or failure.
Bound retries and space them according to the service's rules for transient faults and rate limits. Permission errors or invalid arguments generally need their cause corrected. Repeating an unchanged request in those cases is unlikely to help.
The interface should explain when confirmation is unavailable. Telling the user that a ticket was not created can encourage another request even though the first ticket already exists. Accurate pending-state wording is part of duplicate prevention.
Test the failure intervals
Use a simulated service that can accept a command and discard the response. Restart the worker and inspect the key it reuses. Count the resulting external effects rather than relying only on local success logs.
Add concurrent workers, a late response and changed arguments under an existing key. Revoke a permission while the action waits for approval. Exercise expiry of the provider's idempotency window and confirm which outcomes require manual reconciliation.
The test should explain why one intended operation produces one ticket under those conditions, and where uncertainty remains. It does not establish exactly-once execution across arbitrary distributed systems. A useful handover states the tested assumptions and the operational procedure for unresolved cases.
Review one write integration first
Choose a single state-changing tool for the initial review. Identify the endpoint contract and agree on the operation states and failure tests. Include the process for finding and resolving stuck operations after deployment.
LLM integration work should cover this backend behaviour. The agent runtime guide places the operation within a complete execution flow.
Share the API and an action that must not be duplicated. That allows a bounded integration change to be scoped around retry and reconciliation tests before increasing the agent's autonomy.
Related service
Agent workflows
Connect tools with permissions, state and human review.