mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Automation

Sequential Pipelines Are Killing Your Agent Throughput: Concurrent Execution Patterns That Cut Latency by 3x

How to identify parallelizable agent tasks, orchestrate concurrent execution with promises and DAGs, and handle partial failures without breaking depend...

Source: dev.to
Sequential Pipelines Are Killing Your Agent Throughput: Concurrent Execution Patterns That Cut Latency by 3x

A user abandoned a research report at second 41. The pipeline was still running. Five agents, each waiting for the previous one to finish, even though three of them had no actual dependency on each other. By the time the output landed, the tab was closed.

Sequential orchestration looks reasonable on a whiteboard. Planner decomposes the query. Retriever fetches sources. Analyzer produces findings. Verifier checks claims. Synthesizer merges everything. Five stages, each logically dependent on the last.

In production, it is a serial queue of LLM calls. The retriever cannot start until the planner finishes. The analyzer cannot start until every retrieval completes. The verifier waits on the analyzer. The synthesizer waits on everything. Every agent is idle while the previous one finishes work that has no actual dependency on it.

If market analysis takes 12 seconds and risk assessment takes 10 seconds, sequential execution takes 22 seconds. Parallel execution completes in 12 seconds (the maximum of the two, not the sum). For workflows with independent tasks, parallel execution cuts latency proportionally to the number of concurrent agents.

The Dependency Graph You Are Not Drawing

Most agent pipelines have fewer dependencies than their sequential structure implies. The problem is not knowing which tasks can run in parallel without breaking correctness.

Start by drawing the actual dependency graph:

  • Planner: No dependencies (entry point)
  • Retriever: Depends on planner output (sub-questions)
  • Analyzer: Depends on retriever output (sources)
  • Verifier: Depends on analyzer output (findings)
  • Synthesizer: Depends on verifier output (validated findings)

This looks like a strict chain. But zoom in on the retriever. If the planner produces four sub-questions, the retriever can fetch sources for all four simultaneously. If the analyzer produces findings for each sub-question, those analyses can run in parallel. If the verifier checks claims independently, those checks can run concurrently.

The real dependency graph has three parallelization points:

  1. Retrieval: Fan out across sub-questions
  2. Analysis: Fan out across retrieved sources
  3. Verification: Fan out across claims

A sequential pipeline treats these as single-threaded loops. A concurrent pipeline treats them as parallel map operations.

Orchestration Primitives That Actually Work

You need three primitives to orchestrate concurrent agent execution:

1. Promise.all for Independent Tasks

When multiple agent calls have no dependencies on each other, launch them all and wait for the slowest one:

// Sequential: 40 seconds total
const marketAnalysis = await agent.analyze('market trends');
const riskAssessment = await agent.analyze('risk factors');
const competitorScan = await agent.analyze('competitor landscape');
const regulatoryCheck = await agent.analyze('regulatory environment');

// Concurrent: 12 seconds (slowest agent)
const [marketAnalysis, riskAssessment, competitorScan, regulatoryCheck] = 
  await Promise.all([
    agent.analyze('market trends'),
    agent.analyze('risk factors'),
    agent.analyze('competitor landscape'),
    agent.analyze('regulatory environment')
  ]);

This works when tasks share no state and produce independent outputs. If one agent call fails, the entire batch fails. That is the correct behavior for truly independent work.

2. Promise.allSettled for Partial Failure Tolerance

When you need results from successful agents even if some fail:

const results = await Promise.allSettled([
  agent.analyze('market trends'),
  agent.analyze('risk factors'),
  agent.analyze('competitor landscape'),
  agent.analyze('regulatory environment')
]);

const successful = results
  .filter(r => r.status === 'fulfilled')
  .map(r => r.value);

const failed = results
  .filter(r => r.status === 'rejected')
  .map(r => ({ reason: r.reason }));

// Proceed with partial results or retry only failed tasks

This pattern is critical when agent calls hit rate limits, timeouts, or transient API failures. You get partial results immediately and can retry only the failed subset.

3. DAG Execution for Complex Dependencies

When tasks have mixed dependencies (some parallel, some sequential), model the workflow as a directed acyclic graph:

const dag = {
  planner: { deps: [], fn: () => agent.plan(query) },
  retriever: { deps: ['planner'], fn: (plan) => agent.retrieve(plan.subQuestions) },
  analyzer: { deps: ['retriever'], fn: (sources) => agent.analyze(sources) },
  verifier: { deps: ['analyzer'], fn: (findings) => agent.verify(findings) },
  synthesizer: { deps: ['verifier'], fn: (validated) => agent.synthesize(validated) }
};

async function executeDag(dag) {
  const results = {};
  const completed = new Set();
  
  while (completed.size < Object.keys(dag).length) {
    const ready = Object.entries(dag)
      .filter(([name, node]) => 
        !completed.has(name) && 
        node.deps.every(dep => completed.has(dep))
      );
    
    const batch = await Promise.all(
      ready.map(async ([name, node]) => {
        const depResults = node.deps.map(dep => results[dep]);
        results[name] = await node.fn(...depResults);
        completed.add(name);
      })
    );
  }
  
  return results;
}

This executor runs all tasks with satisfied dependencies in parallel. The planner runs first. Once it completes, the retriever runs. Once the retriever completes, all analyzer calls run concurrently. The DAG shape determines the parallelism automatically.

Observability Changes When Five Agents Run Simultaneously

Sequential pipelines have simple observability. One span per agent call, nested in a linear trace. Concurrent pipelines require different instrumentation.

Trace Structure

Each parallel batch becomes a single parent span with multiple child spans:

research_report (45s → 15s)
├─ planner (3s)
└─ retrieval_batch (12s)
   ├─ retrieve_market (12s)
   ├─ retrieve_risk (8s)
   ├─ retrieve_competitor (10s)
   └─ retrieve_regulatory (7s)
└─ analysis_batch (8s)
   ├─ analyze_market (8s)
   ├─ analyze_risk (6s)
   ├─ analyze_competitor (7s)
   └─ analyze_regulatory (5s)
└─ verification_batch (4s)
   ├─ verify_claim_1 (4s)
   ├─ verify_claim_2 (3s)
   └─ verify_claim_3 (2s)
└─ synthesizer (2s)

The parent span duration is now the maximum of its children, not the sum. This makes waterfall views accurate.

Metrics That Matter

Track these for concurrent execution:

  • Batch utilization: Percentage of agents in a batch that complete within 10% of the slowest agent (high utilization means well-balanced work)
  • Straggler rate: Percentage of batches where one agent takes 2x longer than the median (indicates load imbalance or rate limiting)
  • Partial failure rate: Percentage of batches where at least one agent fails but the batch proceeds (indicates resilience)
  • Concurrency ceiling: Maximum number of simultaneous agent calls before throughput stops improving (indicates rate limit or resource contention)

Error Attribution

When five agents run concurrently and one fails, you need to know which one without replaying the entire batch. Tag each agent call with:

  • Batch ID (groups concurrent work)
  • Task ID (identifies the specific agent call)
  • Dependency path (shows what upstream work this depends on)

This lets you retry a single failed agent call without re-running the entire batch.

Failure Modes You Will Hit

Concurrent execution introduces failure modes that do not exist in sequential pipelines.

Failure ModeCauseMitigation
Rate limit cascadeFive agents hit the same API simultaneously, all get 429sSemaphore to limit concurrent calls per provider
Memory spikeAll agents load large context simultaneouslyStream responses or limit batch size
Partial state corruptionOne agent writes shared state while another reads itImmutable inputs, copy-on-write outputs
DeadlockAgent A waits for B, B waits for A (cycle in DAG)Topological sort before execution
Straggler amplificationOne slow agent blocks an entire batchTimeout per agent, proceed with partial results
Observability explosion5x more spans overwhelm your tracing backendSample traces or aggregate batch spans

The most common failure is rate limit cascade. If your LLM provider allows 10 requests per second and you launch 20 concurrent agent calls, half will fail immediately. Use a semaphore to limit concurrency:

const semaphore = new Semaphore(10); // Max 10 concurrent calls

async function rateLimitedCall(fn) {
  await semaphore.acquire();
  try {
    return await fn();
  } finally {
    semaphore.release();
  }
}

const results = await Promise.all(
  tasks.map(task => rateLimitedCall(() => agent.call(task)))
);

This queues excess calls instead of failing them immediately.

When Sequential Is Still Correct

Concurrent execution is not always faster. Use sequential orchestration when:

  • Agents share mutable state: If one agent modifies a database record and another reads it, sequential execution guarantees consistency
  • Later agents need all prior outputs: If the synthesizer requires every analyzer output (not just any subset), you cannot proceed with partial results
  • Token budgets are tight: Concurrent execution can spike token usage if all agents load large contexts simultaneously
  • Debugging is more important than latency: Sequential traces are easier to read and replay

The decision is not about performance. It is about whether your workflow has actual dependencies or just appears to have them because you wrote it as a loop.

Technical Verdict

Use concurrent execution when:

  • Agent tasks are independent (no shared state, no sequential dependencies)
  • Latency matters more than simplicity (users abandon at 41 seconds)
  • You can tolerate partial failures (some results are better than no results)
  • Your observability stack can handle 5x more spans

Avoid concurrent execution when:

  • Agents modify shared state (databases, files, external APIs with side effects)
  • Later stages require complete outputs from earlier stages (no partial result tolerance)
  • Your rate limits are tight (concurrent calls will cascade into 429s)
  • Debugging is harder than the latency win justifies

The 3x improvement comes from running independent work in parallel instead of in sequence. If your pipeline has no independent work, concurrent execution will not help. Draw the dependency graph first. If it is a straight line, stay sequential. If it has branches, parallelize them.


Tags

agentic-ai orchestration infrastructure

Primary Source

dev.to ↗