mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Dev Tools

AWS Agent Toolkit: Production Plumbing for MCP Servers, Skills, and Plugins

How AWS structures MCP servers, skills, and plugins into a production-ready agent system with authentication boundaries and deployment orchestration.

Source: github.com
AWS Agent Toolkit: Production Plumbing for MCP Servers, Skills, and Plugins

AWS shipped an official agent toolkit that wires MCP servers, skills, and plugins into a single system for production agent deployment. With 2,768 stars and support for Claude Code, Codex, Cursor, and 10+ other coding agents, it’s the first major cloud provider to release a comprehensive, officially-supported agent integration layer.

The toolkit covers service selection, CDK/CloudFormation, serverless, containers, storage, observability, billing, SDK usage, and deployment. It also includes specialized modules for Amazon Bedrock agents, DevSecOps workflows, and data analytics pipelines. This is not a demo. It’s a production-grade system that handles the gap between local agent experimentation and multi-tenant cloud deployment.

Architecture: Three Abstraction Layers

AWS structures the toolkit around three distinct layers:

MCP Servers provide the protocol boundary. They expose AWS service capabilities through the Model Context Protocol, handling request/response serialization, connection management, and protocol-level error handling. Each MCP server maps to a logical domain (core, agents, data-analytics, devsecops).

Skills encapsulate business logic. A skill is a unit of agent capability that combines multiple AWS API calls, state management, and error recovery into a single callable operation. Skills handle the orchestration flow: “deploy a Lambda function” becomes a sequence of IAM role creation, code packaging, function deployment, and permission configuration.

Plugins are the integration surface. They adapt MCP servers and skills to specific agent platforms (Claude Code, Cursor, Codex). Plugins handle platform-specific authentication flows, UI rendering, and command registration.

The boundary between these layers matters. MCP servers are stateless and credential-agnostic. Skills carry state across multiple API calls but don’t know about the calling agent. Plugins handle all platform-specific concerns and credential management.

Authentication and Authorization Model

The toolkit uses a three-tier credential model:

  1. Local development: Agents inherit credentials from the AWS CLI profile. The aws configure agent-toolkit command sets up a dedicated profile with scoped permissions.

  2. Plugin-level authentication: Each plugin maintains its own credential context. When you install aws-core@claude-plugins-official, the plugin requests AWS credentials through the agent platform’s secure input mechanism.

  3. Service-level authorization: Skills use AWS IAM policies to scope permissions. The toolkit ships with least-privilege policy templates for each skill domain.

Credential rotation happens at the plugin layer. When a session expires, the plugin re-prompts for credentials without disrupting the MCP server or skill state. This separation means you can rotate credentials mid-operation without losing agent context.

Audit trails flow through CloudTrail. Every AWS API call made by an agent includes the agent identifier, plugin version, and skill name in the user-agent string. This gives you full traceability from agent action to AWS service invocation.

State Management and Error Recovery

The toolkit handles state persistence through two mechanisms:

Ephemeral state lives in the MCP server process. When an agent asks to “create a Lambda function,” the skill maintains a state machine (role creation, code upload, function deployment) within the server’s memory. If the server crashes, the operation fails cleanly with a rollback.

Durable state lives in AWS services. Skills that require multi-step workflows (like CDK deployments) write checkpoints to S3 or DynamoDB. If an agent disconnects mid-deployment, the next invocation can resume from the last checkpoint.

Error handling follows a fail-fast pattern. Skills don’t retry AWS API calls automatically. Instead, they return structured error responses with remediation hints. The agent decides whether to retry, adjust parameters, or escalate to a human.

Observability and Debugging

The toolkit exposes three observability layers:

LayerMechanismUse Case
ProtocolMCP server logsDebug connection issues, malformed requests
SkillCloudWatch LogsTrace multi-step workflows, API call sequences
ServiceX-Ray tracesProfile AWS service latency, identify bottlenecks

Each MCP server writes structured JSON logs to stdout. Skills emit CloudWatch log groups with a consistent naming pattern: /aws/agent-toolkit/{plugin}/{skill}. X-Ray tracing is opt-in but recommended for production deployments.

The toolkit includes a debugging mode that captures full request/response payloads. Enable it with AWS_AGENT_TOOLKIT_DEBUG=true. This writes sensitive data to logs, so only use it in development.

Deployment Patterns

The toolkit supports three deployment shapes:

Local development: MCP servers run as child processes of the agent platform. The agent spawns a server when you install a plugin and kills it on exit. This is the default for Claude Code and Cursor.

Shared server: Multiple agents connect to a single MCP server instance. This reduces memory overhead and enables cross-agent state sharing. Deploy the server as a systemd service or Docker container.

Serverless: MCP servers run as Lambda functions behind API Gateway. Agents connect over HTTPS instead of stdio. This adds latency (cold start + network) but enables multi-tenant deployments with per-agent billing.

The toolkit includes CloudFormation templates for shared server and serverless deployments. The templates handle IAM roles, VPC configuration, and CloudWatch alarms.

Failure Modes and Mitigations

Common failure scenarios:

Credential expiration mid-operation: Skills checkpoint state before long-running operations. If credentials expire, the agent can resume with fresh credentials.

Rate limiting: Skills respect AWS service quotas and implement exponential backoff. If a quota is exceeded, the skill returns a structured error with the retry-after timestamp.

Partial deployments: CDK and CloudFormation skills use stack rollback by default. If a deployment fails halfway, AWS automatically reverts to the previous state.

Plugin version skew: The toolkit uses semantic versioning. MCP servers reject requests from plugins with incompatible major versions. This prevents protocol mismatches.

Network partitions: MCP servers timeout after 30 seconds of inactivity. If an agent disconnects, the server cleans up resources and terminates.

Code Example: Custom Skill Integration

If you need to extend the toolkit with a custom skill, the pattern looks like this:

from aws_agent_toolkit import Skill, SkillContext
from aws_agent_toolkit.auth import require_permissions

class DeployStaticSite(Skill):
    name = "deploy_static_site"
    description = "Deploy a static site to S3 + CloudFront"
    
    @require_permissions([
        "s3:CreateBucket",
        "s3:PutObject",
        "cloudfront:CreateDistribution"
    ])
    async def execute(self, ctx: SkillContext, site_path: str):
        # Create S3 bucket with versioning
        bucket = await ctx.s3.create_bucket(
            Bucket=f"{ctx.project_name}-{ctx.environment}",
            Versioning={"Status": "Enabled"}
        )
        
        # Upload site files
        for file in ctx.fs.walk(site_path):
            await ctx.s3.upload_file(
                file.path,
                bucket.name,
                file.relative_path
            )
        
        # Create CloudFront distribution
        distribution = await ctx.cloudfront.create_distribution(
            OriginDomainName=bucket.website_endpoint,
            DefaultCacheBehavior={
                "ViewerProtocolPolicy": "redirect-to-https"
            }
        )
        
        return {
            "bucket": bucket.name,
            "distribution_id": distribution.id,
            "url": f"https://{distribution.domain_name}"
        }

The SkillContext object provides authenticated AWS clients, filesystem access, and logging. The @require_permissions decorator validates IAM policies before execution.

Technical Verdict

Use the AWS Agent Toolkit when:

  • You need production-grade agent integration with AWS services
  • You want official support and regular updates from AWS
  • You’re building multi-tenant agent systems with credential isolation
  • You need audit trails and observability for agent actions
  • You’re deploying agents across multiple platforms (Claude, Cursor, Codex)

Avoid it when:

  • You need sub-100ms latency (MCP protocol adds overhead)
  • You’re building single-purpose automation (use AWS SDK directly)
  • You need custom authentication flows (toolkit assumes IAM)
  • You’re working with non-AWS infrastructure
  • You need to support legacy agent platforms without MCP

The toolkit shines in production environments where you need the full stack: authentication, authorization, state management, observability, and deployment orchestration. It’s overkill for simple scripts but essential for multi-agent systems that need to scale.