A Slack AI agent can finish three of five tools and still leave the thread looking broken. Operations leaders and platform engineers see it when a knowledge write succeeds, a ticket create times out, and a blind retry opens a second ticket, or when the external write lands and the Slack reply never does. The requester thinks the agent failed. Support asks why two Intercom tickets point at the same mention. Security asks which write was authoritative.
If you own multi-tool Slack agents for knowledge retrieval, support drafting, or GitHub-aware ops work, use this as a mid-workflow recovery plan. It covers why tool failures are non-atomic, how to checkpoint each step, verify-before-retry, separate mutation truth from Slack reply delivery, suspend ambiguous runs into a dead-letter path, audit fields, failure modes, and a pilot checklist. Use it when a single in-flight run must recover without duplicating side effects. Intake event retries are a different layer; see how to stop Slack event retries from duplicating AI agent actions for claim-store design at the Events API boundary.
Why mid-workflow tool failures break Slack agents
Multi-tool Slack agents do not fail as a single transaction. A run searches Kipwise, drafts a support reply, creates a ticket, and posts a thread update. Each tool call can succeed, fail, time out after dispatch, or return an error after the remote system already applied the change. That is the non-atomic failure pattern described in research on verified tool calls under non-atomic failures: timeouts after dispatch, delayed visibility, and partial state updates cause retry-only agents to duplicate actions even when overall task success rates look fine.
Slack's Agent design guide asks builders to preserve completed work when something goes wrong mid-task, tell the user what succeeded, and offer clear next steps such as retry, modify, or escalate. Leaving the session stuck in a loading state, or discarding three successful steps because the fourth timed out, breaks that contract.
AWS durable-execution guidance on idempotency and retries treats at-least-once retry as safe only when the side effect is idempotent or protected by an idempotency key. A Slack agent that retries create_ticket after an ambiguous timeout without checking whether the ticket already exists will create duplicates under load.
This is distinct from Events API intake retries. Slack may redeliver an app_mention when your endpoint misses the three-second ack window. That problem belongs to durable event claims. Mid-workflow recovery starts after you already accepted the mention and began a multi-tool run.
Checkpoint every tool step in a run ledger
Treat each agent run as a durable workflow with a run ID and ordered steps. Before the first mutating tool, persist:
- Run identity: Slack
team_id,channel,thread_ts, requesting user, and a stablerun_id. - Step plan: the intended tools in order, with a step ID for each.
- Step state:
pending,succeeded,failed,ambiguous, orskipped, plus timestamps. - Side-effect evidence: external IDs returned by tools (ticket ID, PR URL, knowledge page ID, message
ts). - Delivery state: whether the Slack reply for this run was accepted (
pending,sent,failed).
Kipwise Agent is designed as a thread-native work agent: a teammate mentions it once and continues with ordinary replies; applicable actions use the requesting user's permissions; it cites sources it actually opened; it refuses to guess; consequential changes stay explicit. Those behaviors only hold if each tool step leaves recoverable evidence when the process dies mid-run.
Keep the ledger outside the model context window. The model may summarize incorrectly after a crash. Your recovery code must read structured step state, not a regenerated plan.
Verify external postconditions before retrying
Verify external postconditions before retrying after an ambiguous tool timeout; a timeout is not proof the side effect failed. A connector that returns HTTP 504 after 30 seconds may have already created the ticket. The verified tool calls paper shows that verify-before-retry wrappers reduce duplicate actions compared with retry-only baselines under injected non-atomic faults.
Practical verify-before-retry rules for Slack agents:
| Ambiguous tool outcome | Check before retry | If already done | If not done |
|---|---|---|---|
| Create Intercom/helpdesk ticket | Lookup by idempotency key or deterministic external reference | Reuse ticket ID; mark step succeeded | Create once with the same key |
| Write or update Kipwise knowledge page | Read page version or content hash for the intended path | Skip write; reuse page ID | Write once |
| Open GitHub PR or issue | Search by branch name, title marker, or idempotency label | Reuse URL; mark succeeded | Create once |
| Post Slack reply in thread | Lookup message by stored delivery receipt or deterministic client message ID | Skip post; mark delivery sent | Post once |
| Read-only search or metrics fetch | Usually safe to retry | N/A | Retry with backoff |
Generate the idempotency key inside the durable step once, then pass the same key on every attempt, as AWS durable-execution idempotency guidance recommends for side-effecting APIs. If the provider returns a duplicate-request success, treat it as success, not failure.
Do not ask the model "did the ticket get created?" as your only check. The model is not the system of record. The ticket system is.
Separate mutation success from Slack reply delivery
Separate mutation success from Slack reply delivery so recovery can retry a missing reply without replaying a write that already landed. Builder.io's agent-native PR #2212 reports a production failure mode where a Slack correction successfully updated content but left the thread silent. Retrying the whole integration task would risk replaying the mutation. Their fix stages delivery-only payloads, persists receipts and conversation message IDs, and retries only the missing provider reply.
Copy that split into your Slack agent runtime:
- Mutation phase: execute tools that change knowledge, tickets, code, or external systems. Persist each result and external ID before moving on.
- Compose phase: build the user-visible Slack answer from those results and citations.
- Delivery phase: post or stream the reply, then store the Slack message
tsas a delivery receipt.
If delivery fails after mutations succeeded, recovery must not re-enter the mutation phase. It should recompose only if needed, then retry delivery with the same conversation identity. Slack does not expose a generic exactly-once primitive for replies, so delivery remains at-least-once when acceptance is ambiguous before a receipt returns. The durable guarantee you can own is: never rerun a successful mutation to fix a missing reply.
This also matches Slack agent session guidance: your app owns status transitions. A successful write with a failed reply should leave the session recoverable, not stuck forever in processing.
Suspend ambiguous runs into a dead-letter path
Set Slack session status to suspended and park a dead-letter recovery record when replay may duplicate a non-idempotent side effect. Slack's agents.sessions.setStatus method accepts suspended when the agent cannot make progress until the user intervenes, for example when clarification or tool approval is required. Agent sessions docs place suspended in the lifecycle after processing and before the user continues.
A dead-letter queue is the holding place for work that should not auto-retry. AWS defines a DLQ as a queue for messages that cannot be processed because of errors. OpenClaw's agent DLQ guide applies that to chat agents: when a provider times out after tools run, hold the run for review with tool history and request IDs instead of blind replay; when a tool may have changed an external system, pause for human review. Treat that write-up as practitioner guidance about a specific framework, not as a claim about every Slack app.
Minimum DLQ recovery record for a Slack agent:
- Message identity: event or request ID, received time, channel, thread, account.
- Processing state: which steps succeeded, failed, or are ambiguous.
- Failure class: timeout, permission denial, malformed input, unavailable destination, unknown.
- Side-effect evidence: tool-call IDs, external action IDs, delivery attempts.
- Recovery owner: the ops or platform person who can approve replay.
Default policy:
| Situation | Default action |
|---|---|
| Confirmed transient outage, no side effect completed | Bounded retry with backoff |
| Token, routing, or permission misconfiguration | Repair config, then replay once |
| Tool may have changed an external system | Suspend + human review |
| Malformed or out-of-scope input | Reject and retain audit record |
| Unknown failure after a restart | Inspect ledger before any replay |
Slack agent design says to call agents.sessions.setStatus with active or suspended so the app is not stuck thinking indefinitely. Developing an agent tells builders to clear loading state on errors. Do not leave users staring at "Working..." while the run sits in an unowned failure.
Worked example: support draft plus ticket create
Scenario: a CS ops lead mentions the agent in a customer escalation thread. The intended plan is:
- Search Kipwise for the refund policy (read-only).
- Draft a customer-ready reply from help-center evidence.
- Create an Intercom ticket with the draft and thread link.
- Post the draft and ticket link back into Slack.
Failure: step 3 times out after 45 seconds. The Intercom API may or may not have created the ticket. Step 4 never ran. The Slack session is still processing.
Recovery sequence:
- Mark step 3
ambiguousin the run ledger. Do not mark the run failed-and-discarded. - Verify: look up Intercom by the deterministic idempotency key for this
run_id+ step. - If the ticket exists, store its ID, mark step 3
succeeded, compose the Slack reply, and enter delivery-only recovery. - If the ticket does not exist, create once with the same key, then deliver.
- If verification itself is unavailable (Intercom outage), set session status to
suspended, write a DLQ record with side-effect evidence so far, and tell the thread what completed (policy search and draft) plus what is waiting (ticket create). Offer retry, escalate, or cancel. - Never create a second ticket because "the agent looked stuck."
The same pattern applies when step 3 succeeded and step 4 failed: reuse the ticket ID and retry only Slack delivery.
Failure modes to design for
- Process crash after mutation, before ledger write. Persist the intent to mutate (key + step) before the call, then the result after. On restart, verify by key before calling again.
- Model regenerates a different plan after resume. Recovery must follow the ledger, not a newly sampled tool list.
- User clicks stop during processing. Agent sessions require you to handle
agent_session_stopped, stop in-progress work, and move status away fromprocessingyourself. Preserve completed steps; do not auto-replay mutations. - Delivery succeeds twice under ambiguous acceptance. Prefer deterministic client-side message identity where the API allows, and treat a second identical reply as a known residual risk you monitor rather than a reason to replay mutations.
- DLQ becomes an ignored archive. Assign an owner, alert on age, and require the same approval standard for consequential replay as for the original write.
Pilot checklist
- Add a durable run/step ledger keyed by Slack thread and
run_id. - Persist idempotency keys before every mutating tool call.
- Implement verify-before-retry for ticket, knowledge, and GitHub writes.
- Split mutation, compose, and delivery phases with separate state.
- On ambiguous non-idempotent failure, set
agents.sessions.setStatustosuspendedand enqueue a DLQ record. - Post partial-progress messages that name completed steps and next actions.
- Alert on suspended run age, duplicate external IDs detected on verify, and delivery retries.
- Chaos-test in staging: kill the worker after a successful ticket create and before Slack reply; confirm one ticket and one eventual reply.
- Chaos-test an ambiguous timeout with the ticket already created; confirm verify-before-retry reuses it.
- Measure duplicate side effects, suspended-run clear time, and stuck
processingsessions over a two-week pilot.
What to do next
Pick one production Slack agent that creates tickets or writes knowledge today. Add a run ledger and verify-before-retry for its mutating tools before you add more connectors. Then wire suspended plus a DLQ for ambiguous side effects so mid-run timeouts become recoverable operator work instead of silent duplicates. When Search Console access returns, validate whether non-brand queries around Slack AI agent reliability and tool failure recovery are moving; until then, judge success from pilot metrics: duplicate side effects avoided, suspended runs cleared, and stuck processing sessions reduced.
References
- Slack Developer Docs: Agent design (Errors and recovery) - Preserve partial progress mid-task, offer next steps, and clear stuck status.
- Slack Developer Docs: Agent sessions - Session lifecycle including suspended when user intervention is required.
- Slack API: agents.sessions.setStatus - Method reference for processing, active, suspended, and closed.
- Slack Developer Docs: Developing an agent - Builder guidance for clearing loading state on errors.
- OpenClaw: AI agent dead-letter queues - Practitioner guide on DLQ recovery records and pausing replay when tools may have mutated systems.
- arXiv: Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures - Research on verify-before-retry under non-atomic tool failures.
- Builder.io agent-native PR #2212 - Separating successful mutations from Slack reply delivery.
- AWS Durable Execution: Idempotency and retries - Durable-step and idempotency-key patterns for side effects.
- AWS: What is a dead-letter queue? - DLQ definition for unprocessable messages.
- Kipwise Agent - Thread-native, permission-aware agent product behavior that must survive partial tool failures.
- Kipwise: Stop Slack event retries from duplicating AI agent actions - Adjacent intake-layer idempotency for Events API retries.


