mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

AI Agents

AgentCore Runtime Instances: GPU Colocation and Persistent State for Multi-Agent Workflows

How AWS manages multi-agent workflows with GPU colocation, persistent volumes, and filesystem sharing in a three-agent music production pipeline.

Source: aws.amazon.com
AgentCore Runtime Instances: GPU Colocation and Persistent State for Multi-Agent Workflows

AWS just published a detailed walkthrough of AgentCore Runtime Instances, showing how they deploy multi-agent workflows on persistent GPU infrastructure. The example is a three-agent music production pipeline where agents colocate on one EC2 instance, share a filesystem, and hand work to each other over multiple days.

This is not about payment authorization or spending limits (covered in previous AgentCore coverage). This is about the runtime plumbing: how agents share a GPU, how state persists across sessions, and how the orchestration layer decides when to hand work between agents versus spawning parallel execution.

What AgentCore Runtime Instances Solve

Most multi-agent workflows run into three infrastructure problems:

  1. Cold start latency: Spinning up a new container or Lambda for each agent call adds 5-30 seconds per hop.
  2. Filesystem isolation: Agents that need to share intermediate files (audio stems, video frames, model checkpoints) must coordinate through S3 or a shared database.
  3. GPU waste: If three agents each need a GPU for 10 minutes, you either pay for three idle GPUs or serialize the work and wait 30 minutes.

AgentCore Runtime Instances give you a managed EC2 instance with:

  • Persistent EBS volume mounted at /mnt/workspace
  • GPU colocation for multiple agents
  • Multi-day session support (instance stays warm between invocations)
  • Shared filesystem visible to all agents on the instance

The trade-off is cost. You pay for the instance even when agents are idle. For workflows that run frequently or need low-latency handoffs, this beats cold-start overhead. For infrequent batch jobs, Lambda or ECS Fargate is cheaper.

Architecture: Three-Agent Music Production Pipeline

The example pipeline has three agents:

  1. Composer Agent: Generates MIDI sequences using a fine-tuned music transformer model.
  2. Arranger Agent: Takes MIDI, applies instrumentation rules, and renders audio stems.
  3. Mixing Agent: Balances levels, applies EQ, and exports the final stereo mix.

All three agents run on one g5.xlarge instance (1 NVIDIA A10G GPU, 24 GB GPU memory). The orchestration layer decides which agent runs next based on the current state of the workflow.

Filesystem Coordination

Each agent writes to /mnt/workspace:

/mnt/workspace/
  session_abc123/
    01_composer_output.mid
    02_arranger_stems/
      bass.wav
      drums.wav
      synth.wav
    03_final_mix.wav

The orchestration layer passes the session ID to each agent. Agents read from and write to the session directory. No S3 round trips. No database coordination. Just POSIX filesystem semantics.

This works because all agents share the same EBS volume. If you scale to multiple instances, you need a distributed filesystem (EFS, FSx for Lustre) or fall back to S3 for inter-instance coordination.

GPU Sharing

The A10G has 24 GB of memory. The three models are:

  • Composer: 8 GB
  • Arranger: 6 GB
  • Mixing: 4 GB

Total: 18 GB. Fits comfortably on one GPU.

AgentCore does not enforce GPU memory limits per agent. If the composer agent tries to allocate 20 GB, it will OOM and crash. You must size your instance to fit the sum of your agents’ peak memory usage.

The orchestration layer serializes agent execution by default. Only one agent runs at a time. This avoids GPU contention and simplifies state management. If you want parallel execution, you must explicitly configure it and ensure your models fit in GPU memory simultaneously.

Session Persistence and Recovery

Sessions can span multiple days. The instance stays warm. The EBS volume persists. If the instance terminates (spot interruption, hardware failure, manual stop), the volume survives and reattaches to a new instance.

State recovery depends on what you checkpoint:

  • Filesystem state: Automatic. The new instance sees the same /mnt/workspace.
  • In-memory state: Lost. Agents must reload models and reinitialize.
  • Workflow state: Stored in DynamoDB by the orchestration layer. The new instance resumes from the last completed agent.

For long-running creative workflows, this means you can pause work on Friday and resume Monday without reprocessing. For transactional workflows (API-driven agent calls), session persistence is less useful because each invocation is independent.

Orchestration Flow

The orchestration layer is a state machine that tracks:

  • Current session ID
  • Last completed agent
  • Next agent to invoke
  • Retry count and error state

When an agent completes, it writes a status file to /mnt/workspace/session_abc123/.status:

{
  "agent": "composer",
  "status": "success",
  "output_path": "01_composer_output.mid",
  "next_agent": "arranger"
}

The orchestration layer reads this file, validates the output, and invokes the next agent. If the agent crashes or times out, the orchestration layer retries up to three times before marking the session as failed.

This is simpler than a full DAG executor (Airflow, Prefect) but less flexible. You cannot express complex branching logic or conditional execution. The workflow is a linear chain with optional retries.

Security Boundaries

Agents share a filesystem and GPU. There is no process isolation beyond what the OS provides. If one agent writes malicious code to /mnt/workspace and another agent executes it, you have a problem.

Mitigation strategies:

  • Read-only mounts: Mount shared libraries and model weights as read-only. Only the session directory is writable.
  • Sandboxing: Run each agent in a separate container with limited syscall access (seccomp, AppArmor).
  • Input validation: The orchestration layer validates agent outputs before passing them to the next agent.

AgentCore does not enforce these by default. You must configure them in your instance launch template.

Trade-Offs and Failure Modes

DimensionAgentCore Runtime InstancesLambda + S3ECS Fargate
Cold start0-2 seconds (warm instance)5-30 seconds10-60 seconds
Filesystem sharingNative POSIXS3 round tripsEFS or S3
GPU supportNativeNoYes (with GPU-enabled tasks)
Cost (idle time)High (pay for instance)Low (pay per invocation)Medium (pay per task)
Session persistenceMulti-dayNoneTask-scoped
IsolationWeak (shared OS)Strong (separate Lambda)Medium (separate container)

Failure modes:

  • GPU OOM: If agents exceed GPU memory, the CUDA driver kills the process. The orchestration layer retries, but the session state may be corrupted.
  • Disk full: If agents write too much data, the EBS volume fills. The orchestration layer should monitor disk usage and fail gracefully.
  • Spot interruption: If you use spot instances, AWS can terminate with 2 minutes notice. The orchestration layer must checkpoint frequently and resume on a new instance.

Code Snippet: Agent Invocation

Here is how the orchestration layer invokes an agent:

import boto3
import json

bedrock = boto3.client('bedrock-agent-runtime')

def invoke_agent(session_id, agent_name, input_data):
    response = bedrock.invoke_agent(
        agentId='agent-composer-abc123',
        agentAliasId='PROD',
        sessionId=session_id,
        inputText=json.dumps(input_data),
        enableTrace=True
    )
    
    # Stream response chunks
    for event in response['completion']:
        if 'chunk' in event:
            chunk = event['chunk']
            yield chunk['bytes'].decode('utf-8')
        elif 'trace' in event:
            # Log tool calls and intermediate steps
            print(f"Trace: {event['trace']}")

# Invoke composer agent
session_id = "session_abc123"
input_data = {
    "style": "ambient",
    "duration_seconds": 120,
    "output_path": "/mnt/workspace/session_abc123/01_composer_output.mid"
}

for chunk in invoke_agent(session_id, "composer", input_data):
    print(chunk, end='', flush=True)

The enableTrace=True flag logs tool calls, model invocations, and intermediate steps. This is essential for debugging multi-agent workflows where failures can cascade across agents.

Technical Verdict

Use AgentCore Runtime Instances when:

  • Your workflow needs low-latency handoffs between agents (under 2 seconds).
  • Agents share large intermediate files (audio, video, model checkpoints) and S3 round trips are too slow.
  • You need GPU colocation for multiple agents and can size the instance to fit all models.
  • Sessions span multiple hours or days and you want to avoid reprocessing.

Avoid when:

  • Your workflow is infrequent (less than once per hour) and cold start latency is acceptable.
  • Agents are independent and do not share state or files.
  • You need strong isolation between agents (compliance, multi-tenancy).
  • Your budget is tight and you cannot justify paying for idle GPU time.

For the music production use case, AgentCore Runtime Instances make sense. The workflow is interactive, agents share audio files, and sessions can span days. For batch processing or API-driven workflows, Lambda or ECS Fargate is cheaper and simpler.