4.1.5. Amazon Bedrock AgentCore: Deploying and Operating Agents in Production
💡 First Principle: A framework such as Strands Agents decides how an agent thinks. Amazon Bedrock AgentCore decides where it runs, who it is, what it may touch, and how you see what it did. AgentCore is a set of modular services for running agents securely at scale with any framework and any foundation model. You can use the services together or one at a time, so a LangGraph agent on Amazon EKS can still use AgentCore Memory or Gateway.
Section 4.1 opened with the costs every agent carries: state to manage between steps, a larger security surface in external systems, and a ten-step reasoning chain to trace. In production these become long multi-step sessions, isolation between users, credentials for external systems, tool sprawl, and observability. Building each of these yourself on Lambda, ECS, and Step Functions is possible, but it is the undifferentiated work that AgentCore takes over.
The AgentCore services:
| Service | What it does | Scenario cue |
|---|---|---|
| Harness | Managed agent loop: you declare the model, system prompt, and tools, and AgentCore runs orchestration, tool calls, memory, and tracing | "Configuration-based agent, no orchestration code" |
| Runtime | Serverless hosting for agents and tools, one isolated microVM per session | "Deploy a Strands or LangGraph agent to production" |
| Memory | Short-term (within a session) and long-term (across sessions) memory | "Remember preferences between visits" |
| Gateway | Turns OpenAPI, Smithy, and Lambda targets into MCP tools, connects existing MCP servers, and handles inbound and outbound authentication | "Expose existing APIs to many agents as MCP tools" |
| Identity | Agent workload identities, a token vault, and OAuth 2.0 and API key credential providers | "Act on the user's behalf in Slack or Google" |
| Policy | Deterministic Cedar rules evaluated on every tool call that passes through Gateway | "Block refunds over a limit, even under prompt injection" |
| Observability | OpenTelemetry traces, spans, and metrics in Amazon CloudWatch | "See each reasoning step and tool call" |
| Evaluations | LLM-as-a-judge scoring of agent sessions, traces, and spans with built-in or custom evaluators | "Measure agent quality before and after release" |
| Code Interpreter | Sandboxed execution of generated Python, JavaScript, or TypeScript | "Let the agent run analysis code safely" |
| Browser | Isolated cloud browser with Live View and optional session recording | "Fill in a web form that has no API" |
AWS also documents newer services (Optimization for A/B-tested configuration changes, Registry for cataloging agents and tools, and Payments for agent micropayments). Know that they exist; the services in the table carry the design decisions.
AgentCore Runtime: the details that decide scenarios
- Session isolation. Each
runtimeSessionIdgets a dedicated microVM with its own CPU, memory, and filesystem. When the session ends, the microVM is terminated and its memory is sanitized. Give every user conversation its own session ID; sharing one ID shares the environment. - Lifecycle.
idleRuntimeSessionTimeoutdefaults to 900 seconds (15 minutes).maxLifetimedefaults to 28,800 seconds (8 hours), which is also the ceiling on microVMs. The Instances compute type runs agents on AWS managed EC2 capacity in your account and supports multi-day sessions (up to 14 days) and GPU workloads. - Background work. An agent doing asynchronous work reports
HealthyBusyfrom its/pingendpoint so the session counts as active. Advancingtime_of_last_updateon every ping stops the idle timeout from ever firing, so sessions pile up untilmaxLifetime. - Pricing. Consumption-based: CPU billing follows active processing and typically excludes I/O wait, such as waiting for a model response.
- Protocols. HTTP on port 8080 (
/invocations,/ping, and/wsfor WebSocket), MCP on port 8000 (/mcp), A2A on port 9000, and AG-UI. Callers authenticate with SigV4 or OAuth 2.0. - Versions and endpoints. Every update creates an immutable version. The
DEFAULTendpoint follows the latest version; a named endpoint such asPRODstays pinned until you update it. That makes promotion and rollback one call. Running sessions keep the code they started with until they end.
Deploying a Strands agent to Runtime takes a small wrapper around the existing agent:
from strands import Agent
from bedrock_agentcore.runtime import BedrockAgentCoreApp
agent = Agent(system_prompt="You are an order management assistant.")
app = BedrockAgentCoreApp()
@app.entrypoint
def invoke(payload, context):
result = agent(payload.get("prompt", ""))
return {"result": result.message}
app.run() # serves /invocations and /ping on port 8080
Package it one of two ways:
| Option | Direct code deployment (.zip) | Container image (ARM64, in Amazon ECR) |
|---|---|---|
| Package size | Up to 250 MB | Up to 2 GB |
| New sessions per second | About 25 | About 1.6 |
| Patching | AgentCore patches the language runtime | You rebuild from a current base image |
| Best for | Fast iteration, common frameworks | Large or specialized dependencies, existing container pipelines |
AgentCore Memory keeps short-term memory as the events of a session and builds long-term memory through strategies you add to a memory resource: user preferences (choices and styles), semantic (facts and entities), session summaries (a running summary), and episodic (episodes of actions and outcomes, with reflections that learn which approaches worked). Records are organized by namespaces such as /users/{actorId}/preferences/, and several agents can share one memory resource.
AgentCore Gateway and the built-in tools. Gateway puts every tool behind one MCP endpoint and injects each tool's credentials, so the agent never holds them. Its semantic tool selection lets an agent search a catalog of hundreds of tools for the few that fit the task, rather than loading every description into the prompt. Code Interpreter sessions default to 15 minutes and can run up to 8 hours. A custom interpreter in Sandbox network mode can reach Amazon S3 but not the public internet, Public mode opens the internet, and VPC mode reaches private resources. The AWS managed interpreter (aws.codeinterpreter.v1) uses fixed, most-restrictive defaults and does not take your execution role.
Relationship to Amazon Bedrock Agents Classic. AWS recommends AgentCore for new agent development. The migration map:
| Bedrock Agents Classic | AgentCore equivalent |
|---|---|
| Managed orchestration loop | Managed harness |
| Action groups (OpenAPI or function schema plus Lambda) | Gateway tools exposed over MCP |
| Knowledge base attached to the agent | Knowledge base reached through Gateway or a retrieval tool |
| Session and cross-session memory | AgentCore Memory strategies |
| Return of control, AMAZON.UserInput | Inline function tools in the harness |
| Stage-specific prompt overrides, supervisor routing | Code-defined agent on Runtime (not directly replicated in the harness) |
Use the harness unless you need to own the loop. Choose a code-defined agent (Strands, LangGraph, or custom code) on Runtime when you have an existing codebase, stage-level prompt control, or multi-agent routing that the harness cannot yet express.
⚠️ Exam Trap: "Deploy the agent on AgentCore" does not mean "use a Bedrock model." Runtime works with models in or outside Bedrock (Anthropic, OpenAI, Gemini, and others), and the harness supports Bedrock, OpenAI, Gemini, and OpenAI-compatible providers. An option that swaps the team's model for a Bedrock model because AgentCore requires it is wrong.
⚠️ Exam Trap: Isolation is per session, not per runtime. Deploying one runtime per tenant to keep tenants apart adds overhead without adding isolation; a distinct session ID per conversation already provides it.
Reflection Question: A startup opened its AWS account last month. It needs a customer-support agent that remembers each customer's past issues, calls a ticketing REST API described by OpenAPI, and must never close a ticket for a customer other than the signed-in one. Which AgentCore services cover each requirement, and where does the "only your own tickets" rule get enforced?