Retrying MCP Tool Calls Without Duplicating Side Effects

If you ship an AI agent that talks to external services through the Model Context Protocol (MCP), the moment your user sees a “please try again” error, you are holding two risks at once. Either the tool really did fail and a retry is exactly right, or the tool already succeeded and your retry just duplicated a charge, sent a second message, or launched a second cloud job. This guide is about telling those two situations apart and building retry logic that does not quietly double-charge your customers.

The core idea is simple. A network timeout is not the same thing as a failed operation. When an MCP tools/call times out, the server may have already completed the side effect and only lost the response on its way back to you. Treat the request as ambiguous until proven otherwise, and use a durable effect identifier so a second attempt can be reconciled against the first.

Why MCP retries are uniquely dangerous

MCP is built on JSON-RPC 2.0, and that part of the stack already has a working request identity. The trouble is that JSON-RPC’s request ID is built for correlating one request with one response, not for identifying the operation itself. According to the Eunomia research on agent tool retries, three identities are routinely collapsed into one and should not be:

  • Protocol request ID like JSON-RPC id: 37. Used by the transport to pair a response with a request.
  • Attempt ID like attempt-3. Used by your tracing layer to label a specific execution.
  • Effect ID like launch-report-job-2026-09-14 or a UUID. Used to identify the logical side effect the user actually asked for.

Once you separate those, the retry problem gets smaller. You only need one durable identity to survive across attempts, and that identity has to belong to the effect, not to the wire request.

Why “did the call fail” is the wrong question

Most retry loops in agent code answer the wrong question. They ask, “did the network call succeed?” and then retry on timeout, connection reset, or 5xx response. That is fine for read-only tools. It is reckless for anything that creates, sends, or charges.

The MCP specification exposes an idempotentHint on tool annotations, but the same Eunomia analysis notes that the specification treats those annotations as untrusted hints. You cannot assume a third-party MCP server you connect to in production actually behaves idempotently just because its schema says so. Some do, many do not, and a few will lie on purpose.

So the practical rule becomes: every non-read tool is non-idempotent unless you personally verified otherwise, and you build retries that assume a duplicate is the default risk.

A safe retry decision flow for MCP tool calls

Here is the flow I would build into an MCP client today. It is conservative on purpose, because the cost of a duplicate side effect is usually paid by your user.

  1. Classify the tool. Reads and pure lookups can retry freely. Writes and side-effecting calls default to non-idempotent.
  2. Generate an effect ID before the call. A UUIDv4 or a hash of the canonical inputs plus the user’s intent is enough. Attach it as a metadata field inside the tool arguments, or as a wrapper field the MCP server understands.
  3. Set a tight client timeout. Aim for the server’s p95 plus a small buffer, not the maximum time you are willing to wait. A 30-second timeout that succeeds is better than a 5-minute timeout that masks duplicates.
  4. On timeout, do not retry automatically. Send the call to a reconciliation step instead.
  5. Reconcile by polling or by lookup. If the MCP server supports a status or audit endpoint keyed by effect ID, query it. If the prior call committed, you get the prior result back and stop.
  6. Only retry when reconciliation cannot confirm prior completion. If the server truly does not know, and the user-visible cost of a duplicate is acceptable, you may retry once, with the same effect ID, and stop after that.
  7. Surface the rest to the user. When you cannot prove the first attempt failed and a retry could double-bill, tell the user. They get the choice, not the model.

That is the whole framework. The expensive part is step 5, because most off-the-shelf MCP servers today do not expose a status lookup keyed by effect ID. Which brings us to the next decision.

What to do when the MCP server has no status endpoint

This is the common case in 2026. Many MCP servers wrap REST APIs that themselves do not expose idempotency keys, and the MCP layer does not add one. You have three honest options.

First, wrap the tool yourself. Instead of calling the raw MCP server’s create_invoice tool, call your own wrapper that generates an effect ID, persists it, and on retry calls a server-side “list recent invoices by client and amount” lookup to detect a duplicate. This works well for low-volume, high-value actions like payments and ticket creation.

Second, switch to a call-now, fetch-later pattern. MCP has an evolving Tasks primitive, formalized in SEP-1686, that lets a client submit a task with an effect ID and later retrieve its result by ID. The Eunomia write-up notes that for many workflow-style APIs this is closer to how the underlying service already thinks, and it sidesteps the timeout-but-maybe-succeeded ambiguity by giving you a polling handle instead of a single blocking response.

Third, do not use MCP for that tool. For destructive or expensive side effects, a direct API call with a proper idempotency key sent as an HTTP header is often safer than wrapping a non-idempotent MCP tool in your own retry loop. MCP is a great protocol for capability discovery and tool dispatch; it is not always the best place to enforce financial-grade idempotency.

Distinguishing transient blips from real failures

A retry loop that conflates a 503 with a permanent authorization error will burn through rate limits and still fail. For MCP, the practical categories look like this.

  • Transient and safe to retry: TCP reset, 502, 503, 504, MCP server returns a structured retryable error, your client timeout fired but reconciliation says the effect is not committed.
  • Transient but not safe to retry: any timeout on a non-idempotent tool where reconciliation cannot confirm or deny prior completion. Pause and ask the user.
  • Permanent, never retry: 400, 401, 403, 404, MCP tool returns a structured non_retryable error, schema validation failure, missing required arguments.
  • Unknown, treat as non-idempotent: the MCP server returns a 200 with an empty or malformed body, the JSON-RPC layer times out without a structured error code, or the transport disappears mid-response.

The arXiv field study on MCP production deployments points out that the protocol currently lacks structured error semantics, so most clients fall back to inferring retryability from transport-layer signals. That works for HTTP but undercounts real MCP failure modes. Until richer error codes land in the spec, prefer to be pessimistic and route ambiguous cases to reconciliation rather than automatic retry.

A short checklist you can paste into a runbook

  • Every side-effecting MCP tool has a documented effect ID scheme and a reconciliation plan before it ships.
  • Client timeouts are set to p95 plus a buffer, not to “as long as the user might wait.”
  • Idempotency is enforced server-side whenever possible, with the effect ID passed through tool arguments or a custom header that the MCP server forwards.
  • Retries on ambiguous failures cap at one attempt and require user-visible confirmation for higher-risk actions.
  • Logs record effect ID, attempt ID, and MCP request ID as separate fields so a postmortem can tell attempts apart.
  • Off-the-shelf MCP servers are treated as untrusted until their idempotentHint is verified in a sandbox.

FAQ

Do I need to change the MCP server to support idempotency keys? Often, no. Many servers already pass through extra arguments, so adding an idempotency_key to the tool payload is enough. For servers that strip unknown fields, you will need a small server-side change or a wrapper.

Is idempotentHint in the tool annotation enough on its own? No. Treat it as a hint from a potentially untrusted server. The specification says so, and your users’ money should not depend on a string a third-party server returns about itself.

How many retries are safe? One, with the same effect ID, after reconciliation has failed to confirm prior completion. Anything more is a duplicate-inviting loop with extra steps.

What if my tool is genuinely idempotent, like a status read? You can retry freely with normal exponential backoff and do not need an effect ID. Save the heavier machinery for the calls that can hurt someone if they run twice.

Where does the new Tasks primitive fit? Tasks give you a server-side handle keyed by an ID you choose. For long-running, workflow-style MCP tools, that handle is the cleanest reconciliation primitive available today, and it removes the timeout-but-maybe-succeeded ambiguity by replacing the blocking response with a poll.

Sources