Async agents: how to keep work moving without blocking

An engineering guide to separating independent work, pending results, steering, and recovery in agents that operate for hours or days.

Async agentsTool callingOrchestrationDistributed systems
Wasyra AI Systems
Trust, copilots, and enterprise adoption
Published
October 3, 2026
min read
7 min read
Categoría
AI Systems
5gates for async agents
Conceptual diagram of three asynchronous tasks advancing in parallel, preserving state, and converging at a dependency barrier.

Chapter 01

Waiting for one tool should not freeze the entire agent

An agent investigating an incident may start a log query, request a data-warehouse export, and run a test suite. If it treats every tool as a blocking call, the slowest task stops analysis that does not depend on it. If it continues without modeling dependencies, it may write a root cause before receiving the critical data, retry a job that is still running, or act on a stale answer. The problem is not only latency: it is accurately separating work that can advance from decisions that must wait.

On October 2, 2026, OpenAI published a production guide for the GPT-6 family that brings together steering, asynchronous tools, and delegation for tasks lasting hours or days. The publication appeared one Lima calendar day before this article; it does not mean every mechanism launched that day. Its practical value is the design shift: the model can continue independent work while the application executes a slow tool, but the application remains responsible for registering the job, delivering its result, and preserving authorization boundaries.

Chapter 02

The mechanism: launch, register, advance, and synchronize

With a conventional call, the model's turn pauses until the tool output arrives. With an asynchronous tool, the model emits the call and can reason, answer independent parts, or start other work before the result arrives. Execution does not magically move to the model provider: the application starts the job, preserves the relationship between the call and the external process, and returns the output in a continuation. This separation can hide useful waiting time, but it also turns the orchestrator into a small distributed system.

OpenAI's documentation limits this mode to function and custom tools executed by the application; it does not apply to hosted tools or programmatic calls. It also warns that the current multi-agent mode should not combine async tools with parallel tool calls. Those constraints matter because “asynchronous” does not mean “everything concurrent.” The design needs an explicit graph, clear state ownership, and a barrier that waits only when the next action truly depends on pending data.

Chapter 03

1. Model dependencies, not a list of steps

Represent each unit of work with inputs, expected output, effects, timeout, and dependencies. An inventory lookup and a contract read can run together when neither consumes the other; comparing price, currency, and validity depends on both. The barrier belongs at that comparison, not immediately after launching the first job. This graph avoids two extremes: blocking independent work and allowing a conclusion to use incomplete data. It also reveals the critical path that actually determines latency.

Also separate reads from writes. Searches and effect-free validations are often good candidates for early execution; creating, sending, deleting, or approving requires an additional authorization dependency and a fresh-state check. A user's correction can make a result obsolete even when the job completed successfully. Each node should therefore declare not only which data it depends on, but under which version of the goal, scope, and permissions it remains valid.

Chapter 04

2. Use a durable registry for every pending job

A model-readable handle does not replace transport identity. Keep a table that relates the handle, original call_id, external job, conversation, tenant, goal version, state, timestamps, attempts, and argument hash. OpenAI recommends keeping handles unique throughout the conversation, including after a job completes. That rule prevents a late output from being attached to a newer query with the same name and helps detect duplicate or out-of-order deliveries.

The registry must survive worker restarts and disconnects. Useful minimum states are registered, running, completed, failed, cancelled, expired, and outcome_unknown. Store the output or a verifiable reference, but do not pour secrets or sensitive payloads into the prompt. If the result changes a system, add an idempotency key and a reconciliation mechanism. Correct recovery does not mean repeating everything: first ask whether the effect already happened, then resume from the last confirmed state.

Chapter 05

3. Wait only at the barrier and deliver each result once

A wait tool does not accelerate a job; it states that the next decision requires one or more pending results. It should receive an explicit set of handles, resolve only those jobs, and return clear completed, failed, timed-out, or cancelled states. The application may deliver outputs as soon as they become available without waiting for a barrier. When a wait is used, the documentation recommends returning each output on its original call_id first and the wait call's own status afterward, so the model resumes with the evidence available.

Design delivery as at-least-once and consumption as idempotent. Record which conversation version received the result and prevent a retry from producing two decisions. If a tool finishes after the user changes scope, preserve the evidence but classify it as stale until revalidated. Do not confuse “job finished” with “result accepted”: the first is an infrastructure fact; the second requires checking schema, provenance, freshness, and compatibility with the next decision.

Chapter 06

4. Treat steering as a new version, not a rollback

Mid-turn steering can send a correction while a response is running, but OpenAI makes clear that the message is queued: it does not rewrite output already emitted, undo earlier actions, or cancel tools that have started. An acceptance acknowledgment only means the update is queued. A change from “analyze every account” to “only the north region” must therefore increment the goal version, invalidate affected dependencies, and block any write not yet committed; it cannot assume the broader query stopped running.

Keep a log of every steer, its ID, target version, and the point at which it was applied. The pending queue lives on the WebSocket connection and should not be assumed to survive a disconnect; before replaying, compare your own registry with received events. For material actions, introduce a commit point: prepare the effect, revalidate the goal and authorization, and only then execute. Steering changes the future of the flow; compensating for the past requires another explicit, reviewable operation.

Chapter 07

5. Evaluate recovery, not only time saved

An evaluation should compare blocking and asynchronous flows on equivalent tasks. Measure complete acceptance, wall-clock latency, hidden idle time, cost, maximum pending-job count, and time until detecting that a dependency was needed. Add correctness metrics: stale results consumed, duplicates, unnecessary waits, writes after a steer, and decisions made without every required input. An average latency improvement does not offset one material action based on the previous objective.

Test deliberate failures: a job that completes twice, a restarted worker, a timeout followed by late success, malformed output, changed permissions, steering during approval, and a disconnect with an accepted but unapplied update. Verify that the system produces a handoff containing pending jobs, received evidence, uncertain state, and the next safe action. If it uses compaction for long conversations, keep the operational ledger separately: the compacted item can carry context with fewer tokens, but it is opaque and does not replace an auditable log.

Chapter 08

Hypothetical example: vendor onboarding with three waits

This example is hypothetical. An agent receives an onboarding request and launches a tax validation, duplicate search, and sanctions review in parallel. While they run, it checks the form for required fields, prepares questions about missing data, and summarizes internal policy. It does not write the final recommendation because that depends on all three results. The durable registry links each handle to its call_id, vendor, tenant, request version, and external job.

The user corrects the vendor's country during execution. The steer creates a new version: the duplicate search remains valid, but the tax and sanctions checks become stale and are relaunched with new keys. The barrier waits only for those two replacements. If the connection drops, the orchestrator checks the ledger before replaying. The agent produces a draft and evidence; creating the vendor remains behind fresh approval. Concurrency reduces waiting without turning a late correction into an incorrect write.

Chapter 09

The decision: optimize waiting after securing state

Async tools are worth adopting when a task contains slow, independent, observable operations and wall-clock latency changes the workflow's value. They are not appropriate for every short call or as cover for a design without idempotency, cancellation, or ownership. Start with read-only tools, a small cap on pending jobs, and an explicit barrier. Expand concurrency only when you can explain which result informed each decision and recover an interrupted run without duplicating effects.

The five gates—dependencies, registry, barrier, steering, and recovery—turn “keep working while I wait” into an operable contract. The benefit is not keeping the model busy; it is shortening the critical path without weakening consistency or control. If you need to design a long-running agent with tools, permissions, traces, and evaluations, Wasyra's custom agent service can help build that layer. It is an implementation service separate from the product available at agents.wasyra.com.

Written by

Wasyra AI Systems

Trust, copilots, and enterprise adoption

Wasyra AI Systems covers guardrails, suggestion-first modes, and review design so work assistants earn real adoption.

CopilotsTrustB2B AI
More from this author

Series

AI systems that actually reach production

A series on agents, copilots, and guardrails for bringing AI into real work without breaking trust or operations.

Posts in this series

Keep reading

Keep reading