How to Recover from Mid-Workflow Tool Failures in Slack AI Agents

Connected tool sources feeding one search panel with a last-updated stamp on each result

A Slack AI agent can finish three of five tools and still leave the thread looking broken. Operations leaders and platform engineers see it when a knowledge write succeeds, a ticket create times out, and a blind retry opens a second ticket, or when the external write lands and the Slack reply never does. The requester thinks the agent failed. Support asks why two Intercom tickets point at the same mention. Security asks which write was authoritative.

If you own multi-tool Slack agents for knowledge retrieval, support drafting, or GitHub-aware ops work, use this as a mid-workflow recovery plan. It covers why tool failures are non-atomic, how to checkpoint each step, verify-before-retry, separate mutation truth from Slack reply delivery, suspend ambiguous runs into a dead-letter path, audit fields, failure modes, and a pilot checklist. Use it when a single in-flight run must recover without duplicating side effects. Intake event retries are a different layer; see how to stop Slack event retries from duplicating AI agent actions for claim-store design at the Events API boundary.

Why mid-workflow tool failures break Slack agents

Multi-tool Slack agents do not fail as a single transaction. A run searches Kipwise, drafts a support reply, creates a ticket, and posts a thread update. Each tool call can succeed, fail, time out after dispatch, or return an error after the remote system already applied the change. That is the non-atomic failure pattern described in research on verified tool calls under non-atomic failures: timeouts after dispatch, delayed visibility, and partial state updates cause retry-only agents to duplicate actions even when overall task success rates look fine.

Slack's Agent design guide asks builders to preserve completed work when something goes wrong mid-task, tell the user what succeeded, and offer clear next steps such as retry, modify, or escalate. Leaving the session stuck in a loading state, or discarding three successful steps because the fourth timed out, breaks that contract.

AWS durable-execution guidance on idempotency and retries treats at-least-once retry as safe only when the side effect is idempotent or protected by an idempotency key. A Slack agent that retries create_ticket after an ambiguous timeout without checking whether the ticket already exists will create duplicates under load.

This is distinct from Events API intake retries. Slack may redeliver an app_mention when your endpoint misses the three-second ack window. That problem belongs to durable event claims. Mid-workflow recovery starts after you already accepted the mention and began a multi-tool run.

Checkpoint every tool step in a run ledger

Treat each agent run as a durable workflow with a run ID and ordered steps. Before the first mutating tool, persist:

  1. Run identity: Slack team_id, channel, thread_ts, requesting user, and a stable run_id.
  2. Step plan: the intended tools in order, with a step ID for each.
  3. Step state: pending, succeeded, failed, ambiguous, or skipped, plus timestamps.
  4. Side-effect evidence: external IDs returned by tools (ticket ID, PR URL, knowledge page ID, message ts).
  5. Delivery state: whether the Slack reply for this run was accepted (pending, sent, failed).

Kipwise Agent is designed as a thread-native work agent: a teammate mentions it once and continues with ordinary replies; applicable actions use the requesting user's permissions; it cites sources it actually opened; it refuses to guess; consequential changes stay explicit. Those behaviors only hold if each tool step leaves recoverable evidence when the process dies mid-run.

Keep the ledger outside the model context window. The model may summarize incorrectly after a crash. Your recovery code must read structured step state, not a regenerated plan.

Verify external postconditions before retrying

Verify external postconditions before retrying after an ambiguous tool timeout; a timeout is not proof the side effect failed. A connector that returns HTTP 504 after 30 seconds may have already created the ticket. The verified tool calls paper shows that verify-before-retry wrappers reduce duplicate actions compared with retry-only baselines under injected non-atomic faults.

Practical verify-before-retry rules for Slack agents:

Ambiguous tool outcomeCheck before retryIf already doneIf not done
Create Intercom/helpdesk ticketLookup by idempotency key or deterministic external referenceReuse ticket ID; mark step succeededCreate once with the same key
Write or update Kipwise knowledge pageRead page version or content hash for the intended pathSkip write; reuse page IDWrite once
Open GitHub PR or issueSearch by branch name, title marker, or idempotency labelReuse URL; mark succeededCreate once
Post Slack reply in threadLookup message by stored delivery receipt or deterministic client message IDSkip post; mark delivery sentPost once
Read-only search or metrics fetchUsually safe to retryN/ARetry with backoff

Generate the idempotency key inside the durable step once, then pass the same key on every attempt, as AWS durable-execution idempotency guidance recommends for side-effecting APIs. If the provider returns a duplicate-request success, treat it as success, not failure.

Do not ask the model "did the ticket get created?" as your only check. The model is not the system of record. The ticket system is.

Separate mutation success from Slack reply delivery

Separate mutation success from Slack reply delivery so recovery can retry a missing reply without replaying a write that already landed. Builder.io's agent-native PR #2212 reports a production failure mode where a Slack correction successfully updated content but left the thread silent. Retrying the whole integration task would risk replaying the mutation. Their fix stages delivery-only payloads, persists receipts and conversation message IDs, and retries only the missing provider reply.

Copy that split into your Slack agent runtime:

  1. Mutation phase: execute tools that change knowledge, tickets, code, or external systems. Persist each result and external ID before moving on.
  2. Compose phase: build the user-visible Slack answer from those results and citations.
  3. Delivery phase: post or stream the reply, then store the Slack message ts as a delivery receipt.

If delivery fails after mutations succeeded, recovery must not re-enter the mutation phase. It should recompose only if needed, then retry delivery with the same conversation identity. Slack does not expose a generic exactly-once primitive for replies, so delivery remains at-least-once when acceptance is ambiguous before a receipt returns. The durable guarantee you can own is: never rerun a successful mutation to fix a missing reply.

This also matches Slack agent session guidance: your app owns status transitions. A successful write with a failed reply should leave the session recoverable, not stuck forever in processing.

Suspend ambiguous runs into a dead-letter path

Set Slack session status to suspended and park a dead-letter recovery record when replay may duplicate a non-idempotent side effect. Slack's agents.sessions.setStatus method accepts suspended when the agent cannot make progress until the user intervenes, for example when clarification or tool approval is required. Agent sessions docs place suspended in the lifecycle after processing and before the user continues.

A dead-letter queue is the holding place for work that should not auto-retry. AWS defines a DLQ as a queue for messages that cannot be processed because of errors. OpenClaw's agent DLQ guide applies that to chat agents: when a provider times out after tools run, hold the run for review with tool history and request IDs instead of blind replay; when a tool may have changed an external system, pause for human review. Treat that write-up as practitioner guidance about a specific framework, not as a claim about every Slack app.

Minimum DLQ recovery record for a Slack agent:

  1. Message identity: event or request ID, received time, channel, thread, account.
  2. Processing state: which steps succeeded, failed, or are ambiguous.
  3. Failure class: timeout, permission denial, malformed input, unavailable destination, unknown.
  4. Side-effect evidence: tool-call IDs, external action IDs, delivery attempts.
  5. Recovery owner: the ops or platform person who can approve replay.

Default policy:

SituationDefault action
Confirmed transient outage, no side effect completedBounded retry with backoff
Token, routing, or permission misconfigurationRepair config, then replay once
Tool may have changed an external systemSuspend + human review
Malformed or out-of-scope inputReject and retain audit record
Unknown failure after a restartInspect ledger before any replay

Slack agent design says to call agents.sessions.setStatus with active or suspended so the app is not stuck thinking indefinitely. Developing an agent tells builders to clear loading state on errors. Do not leave users staring at "Working..." while the run sits in an unowned failure.

Worked example: support draft plus ticket create

Scenario: a CS ops lead mentions the agent in a customer escalation thread. The intended plan is:

  1. Search Kipwise for the refund policy (read-only).
  2. Draft a customer-ready reply from help-center evidence.
  3. Create an Intercom ticket with the draft and thread link.
  4. Post the draft and ticket link back into Slack.

Failure: step 3 times out after 45 seconds. The Intercom API may or may not have created the ticket. Step 4 never ran. The Slack session is still processing.

Recovery sequence:

  1. Mark step 3 ambiguous in the run ledger. Do not mark the run failed-and-discarded.
  2. Verify: look up Intercom by the deterministic idempotency key for this run_id + step.
  3. If the ticket exists, store its ID, mark step 3 succeeded, compose the Slack reply, and enter delivery-only recovery.
  4. If the ticket does not exist, create once with the same key, then deliver.
  5. If verification itself is unavailable (Intercom outage), set session status to suspended, write a DLQ record with side-effect evidence so far, and tell the thread what completed (policy search and draft) plus what is waiting (ticket create). Offer retry, escalate, or cancel.
  6. Never create a second ticket because "the agent looked stuck."

The same pattern applies when step 3 succeeded and step 4 failed: reuse the ticket ID and retry only Slack delivery.

Failure modes to design for

  1. Process crash after mutation, before ledger write. Persist the intent to mutate (key + step) before the call, then the result after. On restart, verify by key before calling again.
  2. Model regenerates a different plan after resume. Recovery must follow the ledger, not a newly sampled tool list.
  3. User clicks stop during processing. Agent sessions require you to handle agent_session_stopped, stop in-progress work, and move status away from processing yourself. Preserve completed steps; do not auto-replay mutations.
  4. Delivery succeeds twice under ambiguous acceptance. Prefer deterministic client-side message identity where the API allows, and treat a second identical reply as a known residual risk you monitor rather than a reason to replay mutations.
  5. DLQ becomes an ignored archive. Assign an owner, alert on age, and require the same approval standard for consequential replay as for the original write.

Pilot checklist

  1. Add a durable run/step ledger keyed by Slack thread and run_id.
  2. Persist idempotency keys before every mutating tool call.
  3. Implement verify-before-retry for ticket, knowledge, and GitHub writes.
  4. Split mutation, compose, and delivery phases with separate state.
  5. On ambiguous non-idempotent failure, set agents.sessions.setStatus to suspended and enqueue a DLQ record.
  6. Post partial-progress messages that name completed steps and next actions.
  7. Alert on suspended run age, duplicate external IDs detected on verify, and delivery retries.
  8. Chaos-test in staging: kill the worker after a successful ticket create and before Slack reply; confirm one ticket and one eventual reply.
  9. Chaos-test an ambiguous timeout with the ticket already created; confirm verify-before-retry reuses it.
  10. Measure duplicate side effects, suspended-run clear time, and stuck processing sessions over a two-week pilot.

What to do next

Pick one production Slack agent that creates tickets or writes knowledge today. Add a run ledger and verify-before-retry for its mutating tools before you add more connectors. Then wire suspended plus a DLQ for ambiguous side effects so mid-run timeouts become recoverable operator work instead of silent duplicates. When Search Console access returns, validate whether non-brand queries around Slack AI agent reliability and tool failure recovery are moving; until then, judge success from pilot metrics: duplicate side effects avoided, suspended runs cleared, and stuck processing sessions reduced.

References

Want a better team wiki?
Try Kipwise - integrated with your favorite everyday tools