Designing reliable async workflows in GraphN
- Published
- August 12, 2026
- Reading time
- 9 min
An asynchronous run is not a synchronous request with a longer timeout. It is a lifecycle: accept work, give it a durable identity, execute a published definition outside the client connection, and make the outcome inspectable later.
A workflow that classifies one message may fit comfortably inside a normal HTTP request. A workflow that reads files, calls several agents, waits on tools, or fans out across independent analyses may not. Keeping the original connection open makes the caller responsible for mobile sleep, proxy deadlines, deploys, and retry ambiguity.
GraphN's async interface changes that contract. The caller submits input to a published workflow and receives an execution identifier. It can then poll or watch the execution, inspect logs and trace information when available, or request cancellation without keeping the submission connection open.
This article stays at that public boundary: what an application submits, what it can observe, and what it must still design for itself.
When async is the right interface
Use synchronous execution when the result is required immediately and the workflow is expected to finish inside the caller's request window. It is the simpler contract for an interactive lookup or a small deterministic transformation.
Use async when one or more of these conditions apply:
- the workflow can outlive a browser tab, CLI process, webhook deadline, or reverse-proxy timeout;
- it includes multiple model, retrieval, function, connector, or MCP calls;
- independent branches can run concurrently before a join;
- the caller needs an execution ID it can store and inspect later;
- operators need to distinguish accepted, running, failed, and canceled work;
- cancellation matters even after the submission connection is gone.
Async does not make an unsuitable workflow suitable. A non-idempotent tool can still duplicate an external side effect if it is retried. An unbounded loop is still unsafe. A missing output contract is still hard for a downstream service to consume. Async gives the run a durable lifecycle; it does not remove the need to design the workflow.
The public lifecycle in one diagram
An application needs only the public workflow and execution interfaces. It submits a published workflow, stores the returned identifier, and reads the execution until it reaches a terminal state.
An acceptance response means the request entered the asynchronous lifecycle, not that any workflow step succeeded. Persist the execution ID before navigating away or acknowledging the business event that initiated the run.
If submission is rejected explicitly, handle that error before polling. If the network fails before the client can tell whether the request was accepted, treat the outcome as ambiguous rather than assuming that no work exists.
Retry safety is an application design problem
The dangerous moment in any asynchronous API is the gap between server acceptance and client receipt. A network failure can leave the caller unsure whether work started. Blindly submitting again can create duplicate work.
Define the client's retry policy before production:
- Store the execution ID beside the business record that initiated the run.
- After an ambiguous submit, check the business record and available execution history before creating replacement work.
- Give every side-effecting tool call its own stable business key.
- Never assume that making submission safer also makes downstream actions safe to repeat.
If a workflow step charges a card, creates a ticket, or sends a message, the destination system should reject or safely coalesce duplicate business operations. That control belongs at the side-effect boundary even when the surrounding workflow is asynchronous.
Publish before you submit
A production run should not discover halfway through that an agent prompt, function body, or MCP configuration changed after submission. Publishing is the boundary that prevents this. As described in the workflow publishing guide, a published version captures the workflow DSL and linked resource snapshots.
Applications invoke that published version through the GraphN CLI or workspace-scoped API. They send workflow input rather than implementation configuration. The execution record then gives the application a stable place to inspect status and result information.
Dependencies make progress understandable
The published workflow DSL makes ordering visible through two sources:
- explicit
afteredges; and - references such as
${ steps.lookup.output }, which imply that the consumer waits forlookup.
Steps at the same dependency level can run concurrently. A join runs only after its prerequisites complete.
yaml
steps:
policy_review:
call: agent
agent: Policy_Reviewer
input_template: "${ input.message }"
order_lookup:
call: mcp_tool
server: Store_Tools
tool: get_order_status
input:
ticket_id: "${ input.ticket_id }"
decide:
call: agent
agent: Resolution_Agent
after: [policy_review, order_lookup]
input_template: |
Policy: ${ steps.policy_review.output }
Order: ${ steps.order_lookup.output }
output:
resolution: "${ steps.decide.output }"policy_review and order_lookup have no dependency on each other, so they are ready together. decide is in the next group because it names both predecessors. The top-level graph must remain acyclic; iteration belongs in bounded for_each, while, or compound steps rather than a cycle between normal nodes.Parallel readiness does not make partial success equivalent to overall success. If a required branch fails, the dependent join cannot safely behave as though every input exists. Define a top-level output so downstream applications know which result to consume, and inspect available step or trace details when a run fails.
Status, failure, and cancellation behavior
Scroll horizontally to compare
| State you may observe | What it establishes | What the caller should do |
|---|---|---|
pending | The execution record exists but workflow work has not started. | Continue polling; do not treat it as success. |
running | Workflow work has started. | Continue watching; inspect progress or logs if needed. |
completed | The terminal output was recorded. | Validate and consume output. |
failed | Workflow execution reached an error. | Read the error, failed-step context, and available trace before deciding whether to retry. |
canceled | Cancellation was recorded for the execution. | Stop waiting for output and verify any external side effects separately. |
Treat terminal values explicitly rather than collapsing every non-running state into success or failure. Cancellation is a request to stop remaining workflow work, not a database transaction rollback. A tool call that already committed in an external system may remain committed. Design compensating actions where the business process requires them.
Run and watch from the CLI
Publish the compact order-lookup workflow from the execution-model article, configure the CLI's workspace and API endpoint, then submit JSON that matches its published input schema:
bash
graphn wf publish <workflow-id> -m "Enable production async run"
graphn wf run <workflow-id> \
--mode async \
--input '{"message":"Where is order ORD-4821?","order_id":"ORD-4821"}'The submission returns an execution identifier. Copy it and watch the execution:
bash
graphn exec get <execution-id> --watch--watch polls rather than holding the submission request open. It prints status changes and returns the terminal JSON, including output on success or error information on failure. The example workflow was validated end to end: the explicit lookup completed before the answer step and the final execution completed successfully.You can stop the local watch with
Ctrl+C without canceling the server-side execution; run graphn exec get <execution-id> later to resume inspection. The developer CLI guide covers authentication, workspace selection, and API configuration.What asynchronous execution does not imply
The public GraphN contract is an execution record: identity, status, output or error, logs, and trace detail when available. Do not infer arbitrary user-addressable checkpoints, rewind, branch-from-history, or time-travel semantics from the word “async.” If your application requires human interruption at arbitrary graph boundaries or replay from a selected checkpoint, verify that exact public behavior rather than treating async submission as a substitute.
Production checklist
Before switching a workflow to async:
- Publish the workflow and every linked resource; do not depend on draft edits.
- Define and validate a stable input and output schema.
- Choose async because the lifecycle can outlive the request, not merely because the workflow is slow today.
- Define how the client handles an ambiguous submission result.
- Make side-effecting tools idempotent with domain-level keys.
- Bound loops and fan-out; use explicit dependencies where intent is not obvious.
- Set realistic timeouts and decide which external errors are safe to retry.
- Treat
202 Acceptedas the beginning of monitoring, never as success. - Persist the execution ID beside the application record that caused the run.
- Handle completed, failed, and canceled outcomes explicitly.
- Inspect failed-step context and traces before automatic resubmission.
- Define compensation for external effects that cancellation cannot undo.
- Test an ambiguous submission timeout, a tool timeout, a malformed response, and cancellation during a side effect.
Async execution is most useful when the lifecycle is part of the application design. The execution ID becomes a join key between the initiating business event, the GraphN execution record, and the final trace. That makes a long-running agent workflow something a client can submit, leave, revisit, and reason about—without pretending that acceptance is completion or that durable scheduling is arbitrary time travel. The AI workflow platform overview places this async path in the full compose, publish, run, and inspect lifecycle.
Related production-agent engineering articles
- Publishing GraphN workflows: drafts, resources, and production versions — understand the public publishing boundary before this async path begins.
- Three trust boundaries for production agents — decide whether each capability should use a function, MCP tool, or grounded retrieval resource.
Primary sources
- GraphN AI workflow platform
Defines publishing, synchronous and asynchronous execution, execution identifiers, cancellation, logs, and inspection.
- Developing with GraphN agents
Documents CLI setup, published workflow runs, async mode, and the execution watch command used in this article.
- GraphN API reference
Documents workspace authentication, published workflow snapshots, execution records, outputs, errors, traces, and HTTP error behavior.
- GraphN FAQ: sync and async runs
Defines the public distinction between blocking synchronous runs and asynchronous runs that return an execution identifier.
- GraphN workflow DSL
Documents step dependencies, expression-derived ordering, parallel fan-out and join, bounded loops, guards, and workflow outputs.