mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Automation

The Crash Test Nobody Runs: What Exactly-Once Semantics Mean for AI Workflow Engines Like n8n

How visual workflow tools handle idempotency, retries, and partial failures when orchestrating AI agent steps and external API calls.

Source: dev.to
The Crash Test Nobody Runs: What Exactly-Once Semantics Mean for AI Workflow Engines Like n8n

Visual workflow engines like n8n show green checkmarks when an execution finishes. They do not show whether the side effect (the email, the charge, the database row) happened zero times, once, or twice. That gap between execution state and external effect is where exactly-once semantics collapse.

A developer building n8n workflow templates in public documented the failure modes that surface when you chain LLM calls, API integrations, and stateful operations in a low-code automation platform. The core insight: green execution logs tell you the workflow ran, not that the external action happened exactly once.

Two Crash Shapes

There are two distinct failure modes when a workflow node crashes mid-execution. Only one is reproducible from inside n8n.

Accepted and known, receipt lost. The HTTP node returns, you hold the provider response, and then the process dies before you persist “done”. A Code node that throws after the HTTP call reproduces this exactly. It is deterministic. You already observed acceptance, so you can reconcile from the provider ID if you logged it anywhere durable.

Submitted, response never observed. The request reached the provider and the side effect committed, but your process never saw the response. Connection reset after commit, timeout after the write landed, process killed between send and receive. This is the case that makes exactly-once impossible and forces an Unknown or manual review state.

You cannot reproduce the second case reliably from inside n8n. Whether the request already left the socket when you kill the process is a race, not a control. This is the boundary where testing stops and operational monitoring begins.

What n8n Persists Between Steps

n8n stores execution state in a relational database (Postgres or SQLite). Each workflow execution gets a record. Each node execution gets a record. The engine writes these records after the node completes.

The persistence model:

  • Execution record: workflow ID, start time, end time, status (success, error, waiting)
  • Node execution record: node ID, input data, output data, error details
  • Waiting executions: for workflows with webhook triggers or wait nodes, a separate table holds resumption state

When a workflow crashes mid-execution, the database holds partial state. The execution record exists but may show “running” indefinitely. Node records exist only for completed nodes. Nodes that were in-flight when the crash happened leave no trace.

n8n does not write a node record until the node finishes. If the process dies while an HTTP node is waiting for a response, there is no record of the request being sent. The workflow engine has no way to know whether the external system received the request, processed it, or committed the side effect.

Idempotency Tokens and Deduplication

n8n does not provide built-in idempotency token management. You have to implement it yourself in the workflow.

A typical pattern:

  1. Generate a unique execution ID at the start of the workflow (use {{ $execution.id }} or a UUID node)
  2. Pass that ID as an idempotency key to external APIs that support it (Stripe, Twilio, most payment processors)
  3. Store the execution ID and provider response in a durable store (Postgres, Redis, Airtable)
  4. Before making the external call, check the store to see if this execution ID already completed

The problem: if the workflow crashes between step 2 (making the call) and step 3 (storing the result), you have no record of the call. On retry, the idempotency key prevents a duplicate charge, but you still do not know the outcome. The execution shows as failed, but the side effect succeeded.

This is where exactly-once delivery diverges from exactly-once processing. Message queues like Kafka and RabbitMQ provide exactly-once delivery guarantees by coordinating acknowledgment with the broker. Visual workflow engines do not have that coordination layer. They execute nodes sequentially and persist state after completion. The gap between execution and persistence is where duplicates or lost effects hide.

Retry Behavior and Side Effect Replay

When an n8n workflow fails and you retry it, the entire workflow runs again from the start. There is no automatic checkpoint-and-resume for individual nodes.

What gets replayed:

  • All HTTP calls
  • All LLM invocations (OpenAI, Anthropic, etc.)
  • All database writes
  • All webhook sends

What does not get replayed:

  • Nothing. The retry is a full re-execution.

This is different from durable execution engines like Temporal or Restate, which checkpoint after each step and replay only the failed portion. In n8n, retry means re-run everything.

The implications for AI workflows:

  • If a workflow calls GPT-4 three times and fails on the fourth call, retrying the workflow calls GPT-4 three more times. You pay for six LLM invocations.
  • If a workflow writes to a database and then fails on an HTTP call, retrying the workflow writes to the database again. You need application-level deduplication (unique constraints, upserts, idempotency checks).
  • If a workflow sends a webhook notification and then fails on a logging step, retrying the workflow sends the webhook again. The downstream system receives duplicate events.

State Management Primitives

n8n provides a few primitives for managing state across workflow executions:

PrimitiveScopeDurabilityUse Case
Static dataWorkflow definitionPersisted with workflowConfiguration, API keys, templates
Execution dataSingle executionPersisted after node completionPassing data between nodes in one run
Workflow variablesWorkflow instancePersisted in databaseCounters, flags, last-run timestamps
External store (HTTP node)Cross-executionDepends on external systemDeduplication, reconciliation, audit log

Workflow variables are the closest thing to durable state. You can read and write them from any node using the $vars object. They persist across executions. But they are not transactional. If you increment a counter and the workflow crashes before the next node, the counter still incremented. On retry, it increments again.

External stores (Postgres, Redis, Airtable) give you transactional guarantees if you use them correctly. Write the execution ID and provider response in a single transaction. Check the store before making the external call. This moves the exactly-once boundary from the workflow engine to the external system, where you have better control.

Observability Gaps

n8n’s execution list shows:

  • Workflow name
  • Trigger type
  • Start time
  • End time
  • Status (success, error, waiting)
  • Error message (if failed)

It does not show:

  • Which external systems were called
  • Which calls succeeded and which failed
  • Whether side effects committed
  • Whether retries are safe

You have to instrument this yourself. A common pattern:

// In a Code node before each external call
const executionId = $execution.id;
const nodeId = $node.name;
const timestamp = new Date().toISOString();

// Log the intent
await fetch('https://your-logging-endpoint.com/log', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    executionId,
    nodeId,
    timestamp,
    action: 'intent',
    target: 'stripe.charges.create',
    idempotencyKey: executionId
  })
});

// Make the external call
const response = await fetch('https://api.stripe.com/v1/charges', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${$credentials.stripe.apiKey}`,
    'Idempotency-Key': executionId
  },
  body: new URLSearchParams({
    amount: 1000,
    currency: 'usd',
    source: 'tok_visa'
  })
});

const result = await response.json();

// Log the outcome
await fetch('https://your-logging-endpoint.com/log', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    executionId,
    nodeId,
    timestamp: new Date().toISOString(),
    action: 'outcome',
    target: 'stripe.charges.create',
    idempotencyKey: executionId,
    providerId: result.id,
    status: result.status
  })
});

return { result };

This gives you an external audit log. If the workflow crashes between the intent log and the outcome log, you know the call was attempted but the result is unknown. You can query the provider API using the idempotency key to reconcile.

Deployment Shape and Failure Modes

n8n runs as a Node.js process. You can deploy it as:

  • A single container (Docker, Kubernetes pod)
  • A queue-mode setup (separate webhook and worker processes)
  • A scaled worker pool (multiple worker processes, shared database)

In single-container mode, a process crash loses all in-flight executions. The database holds partial state. Restarting the container does not resume in-flight workflows. You have to manually retry or reconcile.

In queue mode, webhook processes receive triggers and write them to a queue (Redis or RabbitMQ). Worker processes pull from the queue and execute workflows. A worker crash loses the in-flight execution, but the queue message can be retried. This gives you at-least-once delivery at the workflow level, but not at the node level.

In scaled worker mode, multiple workers share a database. A worker crash does not affect other workers. But the crashed execution still shows as “running” in the database. You need a separate process to detect stale executions and mark them as failed or retry them.

Common failure modes:

  • Worker OOM: Long-running workflows with large data payloads exhaust memory. The process crashes. In-flight executions are lost.
  • Database connection timeout: High concurrency exhausts database connections. Workers cannot persist state. Executions complete but leave no record.
  • External API timeout: An HTTP node waits indefinitely for a response. The workflow hangs. No error, no completion, no retry.
  • Webhook delivery failure: A webhook trigger fires, but the workflow crashes before processing. The webhook sender does not retry. The event is lost.

When Exactly-Once Is Actually Possible

Exactly-once semantics require coordination between the workflow engine and the external system. The only way to achieve it:

  1. The external system supports idempotency keys
  2. You generate a stable key (execution ID, workflow ID + timestamp, UUID)
  3. You pass the key with every request
  4. The external system deduplicates based on the key
  5. You store the key and provider response in a durable store
  6. You check the store before making the request

This moves the exactly-once guarantee from the workflow engine to the external system. The workflow engine provides at-least-once execution. The external system provides deduplication. Together, they give you exactly-once side effects.

For systems that do not support idempotency keys (legacy APIs, databases without upsert, webhooks without deduplication), exactly-once is not possible. You have to choose:

  • At-least-once (accept duplicates, handle them downstream)
  • At-most-once (accept lost events, handle gaps downstream)
  • Manual reconciliation (detect duplicates and gaps, fix them manually)

Technical Verdict

Use n8n for AI workflows when:

  • You can instrument external calls with idempotency keys
  • You can tolerate full workflow retries (re-running all LLM calls, all API calls)
  • You can build external observability (logging, reconciliation, audit trails)
  • You understand the gap between execution state and side effect state

Avoid n8n when:

  • You need guaranteed exactly-once processing without external coordination
  • You cannot afford to replay expensive operations (LLM calls, video processing)
  • You need automatic checkpoint-and-resume for long-running workflows
  • You need transactional guarantees across multiple external systems

For those cases, use a durable execution engine (Temporal, Restate, Inngest) or a message queue with exactly-once delivery (Kafka with transactions, RabbitMQ with publisher confirms and consumer acknowledgments).

The crash test nobody runs is the one where you kill the process between sending the request and receiving the response. That is the boundary where exactly-once semantics break. n8n cannot help you there. You have to build it yourself.


Tags

agentic-ai orchestration infrastructure

Primary Source

dev.to ↗