mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

AI Agents

Ambient Agents on AWS: Event-Driven Triggers, SQS Routing, and the ask_human Tool

How AWS builds ambient agents with SQS triggers, Lambda execution, DynamoDB state, and a single ask_human tool for human-in-the-loop approval.

Source: aws.amazon.com
Ambient Agents on AWS: Event-Driven Triggers, SQS Routing, and the ask_human Tool

Most agent tutorials start with a chat prompt. AWS just published a walkthrough that starts with an S3 upload, a cron schedule, or an alert. The difference is not cosmetic. Event-driven agents need different orchestration plumbing: message routing, state persistence across pauses, and a clean boundary for human approval.

Amazon Bedrock AgentCore now supports ambient agents that respond to signals instead of waiting for user input. The architecture uses SQS for event ingestion, Lambda for execution, DynamoDB for state, and a single ask_human tool that pauses the agent until a human reviews the decision. This is not a chatbot with extra triggers. It is a different control flow.

Why Event-Driven Agents Need Different Plumbing

Chat-based agents run inside a request-response cycle. The user sends a message, the agent thinks, the agent replies. State lives in memory or a session store. The conversation ends when the user closes the window.

Ambient agents have no conversation. An S3 file lands, an SQS message arrives, a CloudWatch alarm fires. The agent wakes up, runs tools, makes decisions, and may need to pause for approval. When the human responds hours later, the agent resumes from the exact tool call that triggered the pause.

This requires:

  • Durable state: The agent’s reasoning chain, tool outputs, and pending decisions must survive Lambda cold starts.
  • Message routing: Multiple event sources (S3, EventBridge, SNS) must map to the correct agent workflow.
  • Human-in-the-loop boundary: The agent must serialize its state, send a notification, and wait for approval without blocking the Lambda function.

AWS solves this with SQS as the event router, DynamoDB as the state store, and a single ask_human tool that externalizes the approval step.

Architecture: SQS to Lambda to DynamoDB

The stack has four layers:

  1. Event sources: S3 bucket notifications, EventBridge schedules, CloudWatch alarms, or SNS topics.
  2. SQS queue: All events land here. The queue decouples event producers from agent execution.
  3. Lambda function: Polls SQS, invokes the Bedrock AgentCore runtime, executes tools, and writes state to DynamoDB.
  4. Jobs page: A web UI where humans review pending decisions and approve or reject tool calls.

When an event arrives, SQS triggers the Lambda function. The function loads the agent’s state from DynamoDB (if resuming) or starts fresh (if new). The agent runs until it hits the ask_human tool, at which point the Lambda function writes the paused state to DynamoDB and exits. A notification goes to the Jobs page. When the human approves, a second SQS message triggers the Lambda function again, and the agent resumes.

Message Routing Logic

SQS does not route by content. Every message goes to the same queue. The Lambda function inspects the message body to determine which agent workflow to invoke. The blog post does not specify the routing logic, but a typical pattern looks like this:

def lambda_handler(event, context):
    for record in event['Records']:
        body = json.loads(record['body'])
        
        if 'eventSource' in body and body['eventSource'] == 'aws:s3':
            agent_id = 'document-processor'
        elif 'source' in body and body['source'] == 'aws.events':
            agent_id = 'scheduled-report'
        elif 'AlarmName' in body:
            agent_id = 'alert-responder'
        else:
            agent_id = 'default-agent'
        
        state = load_state(agent_id, body.get('execution_id'))
        result = bedrock_agent_runtime.invoke(agent_id, state, body)
        
        if result['status'] == 'WAITING_FOR_HUMAN':
            save_state(agent_id, result['execution_id'], result['state'])
            send_notification(result['pending_decision'])
        else:
            save_state(agent_id, result['execution_id'], result['state'])

If multiple events arrive simultaneously, Lambda scales horizontally. Each invocation processes one SQS message. If two events target the same agent workflow but different execution IDs, they run in parallel. If they share the same execution ID (a resume after human approval), DynamoDB optimistic locking prevents race conditions.

The ask_human Tool Boundary

The ask_human tool is the only tool that pauses execution. Every other tool (read S3, query database, send email) runs synchronously inside the Lambda invocation. When the agent calls ask_human, the tool returns a special status code. The Lambda function catches this, serializes the agent’s state, writes it to DynamoDB, and exits.

The state includes:

  • The agent’s reasoning chain (all previous tool calls and LLM responses).
  • The pending decision (what the agent wants to do and why).
  • The execution ID (a unique identifier for this workflow instance).

The Jobs page polls DynamoDB for rows with status = WAITING_FOR_HUMAN. When a human approves, the page writes status = APPROVED and sends a new SQS message with the execution ID. The Lambda function loads the state, injects the approval as a tool response, and the agent continues.

What Happens on Rejection

If the human rejects, the page writes status = REJECTED and optionally includes a reason. The Lambda function loads the state, injects the rejection, and the agent decides what to do next. The blog post does not show rejection handling, but typical patterns include:

  • Retry with a modified plan.
  • Escalate to a different human.
  • Abort the workflow and log the failure.

The agent’s prompt must include instructions for handling rejections. Without explicit guidance, the LLM may hallucinate a retry or ignore the rejection.

Preventing Infinite Loops

Event-driven agents can trigger themselves. An agent writes a file to S3, which generates an S3 event, which triggers the same agent again. Without safeguards, this creates an infinite loop.

AWS does not document loop prevention in the blog post, but standard patterns include:

  • Event filtering: S3 notifications can filter by prefix or suffix. If the agent writes to s3://bucket/output/, configure the trigger to ignore that prefix.
  • Execution depth limit: Store a depth counter in DynamoDB. Increment it on each invocation. Abort if depth exceeds a threshold.
  • Idempotency keys: Include a unique key in the event payload. Store processed keys in DynamoDB. Skip events with duplicate keys.

The Lambda function should enforce these checks before invoking the agent.

State Management Trade-Offs

DynamoDB is not the only option for state storage. Here is how the alternatives compare:

StorageLatencyCost per 1M opsMax item sizeQuery flexibilityBest for
DynamoDB5-10ms$1.25 (on-demand)400 KBKey-value onlyHigh-throughput, simple queries
S350-100ms$0.0055 TBNoneLarge state objects, infrequent access
RDS Postgres10-20ms$0.20 (Aurora)1 GBFull SQLComplex queries, relational data
ElastiCache1-2ms$0.02 per hour512 MBKey-value onlySub-10ms latency, ephemeral state

DynamoDB wins for most ambient agents because state size is small (a few KB of JSON), access is by execution ID (a simple key lookup), and cost scales with usage. S3 makes sense if the agent generates large artifacts (PDFs, images) that need to persist. RDS is overkill unless you need to join state across multiple agents or query historical executions.

Observability Gaps

The blog post does not cover observability. Here is what you need to instrument:

  • Lambda duration per agent invocation: Track how long each agent runs. If duration spikes, the agent may be stuck in a reasoning loop.
  • Tool call latency: Measure time spent in each tool. Slow tools (database queries, API calls) block the agent.
  • Human approval time: Track time from ask_human to approval. Long delays indicate bottlenecks in the review queue.
  • DynamoDB throttling: Monitor ConsumedReadCapacityUnits and ConsumedWriteCapacityUnits. If you hit limits, switch to provisioned capacity or batch writes.
  • SQS queue depth: If the queue grows, Lambda is not scaling fast enough or agents are taking too long.

AWS X-Ray can trace requests across SQS, Lambda, and DynamoDB, but it does not capture agent reasoning. You need custom logging inside the Lambda function to record tool calls, LLM responses, and state transitions.

Deployment Shape

The blog post assumes a single Lambda function handles all agent workflows. This works for prototypes but creates coupling in production. A better shape:

  • One Lambda per agent workflow: Each agent gets its own function, queue, and DynamoDB table. This isolates failures and simplifies IAM policies.
  • Shared runtime layer: Package the Bedrock AgentCore SDK and common tools in a Lambda layer. All functions reference the same layer.
  • Centralized Jobs page: A single web app queries all DynamoDB tables and aggregates pending decisions.

This shape costs more (multiple Lambda functions, multiple queues) but reduces blast radius. If one agent breaks, the others keep running.

When to Use This Pattern

Event-driven agents make sense when:

  • The trigger is not a user prompt (file upload, schedule, alert).
  • The agent needs to pause for human approval before taking action.
  • The workflow spans minutes or hours, not seconds.
  • You need audit logs of every decision and approval.

Avoid this pattern when:

  • The agent must respond in real time (use synchronous Lambda or API Gateway).
  • The workflow is purely automated with no human-in-the-loop (skip the ask_human tool and run end-to-end).
  • State is too large for DynamoDB (use S3 or a database).

Technical Verdict

This architecture is production-ready for ambient agents that need human oversight. The SQS-to-Lambda-to-DynamoDB stack is standard AWS plumbing, and the ask_human tool is a clean abstraction for pausing execution. The main risk is infinite loops from cascading events. Add event filtering, depth limits, and idempotency checks before deploying.

The missing pieces are observability (no built-in tracing for agent reasoning) and rejection handling (the blog post only shows approval). You will need custom logging and explicit prompt instructions for what to do when humans say no.

If you already run agents on Bedrock and need to move from chat-based to event-driven, this is the path. If you are starting fresh, consider whether you need the complexity. A simple Lambda function with a few API calls may be enough.