2.1.4. Agentic AI: Agents, Tools, Memory, and MCP
First Principle: A foundation model on its own is stateless and can only produce text from what it is given. An agent wraps that model with a goal, tools it can call, and memory it can consult, and then runs in a loop until the goal is met.
Generative AI answers the prompt in front of it. Agentic AI takes a goal, such as "sort out this customer's late delivery," and works out the steps itself: look up the shipment, create a replacement order, and email the customer. AWS describes agentic AI as an autonomous system that can act independently to achieve pre-determined goals, cycling through perceive, reason, act, and learn.
The building blocks of an agent:
- A model as the reasoning engine: The foundation model reads the goal and the results so far, then decides what to do next. This is why agentic AI builds on generative AI rather than replacing it.
- Tool usage: A tool is an external capability the agent may call, such as an API, an AWS Lambda function, a database query, a code sandbox, or a web browser. The model asks for a tool call, the agent runs it, and the result goes back to the model. Tools give an agent live data and the ability to act. Order status, stock levels, and today's exchange rate belong in tools, not in fine-tuning, because they change after training.
- Memory management: Models forget everything between requests, so memory is stored outside the model and supplied when it is needed.
- Short-term memory keeps the turn-by-turn history of the current session. It is what lets the agent understand that "what about tomorrow?" still refers to Seattle's weather.
- Long-term memory extracts and keeps insights across sessions, such as user preferences, key facts, and session summaries. A traveler who mentioned a window-seat preference months ago can be offered one without asking again. Memory stores can also be shared between agents.
- Workflow orchestration: Something has to decide the order of steps. That can be the model itself (flexible, but less predictable) or a developer-defined workflow (predictable, but fixed). Choose predictability when a process must run the same way every time, such as a compliance filing.
Model Context Protocol (MCP): connecting agents to external systems
Every tool used to need custom integration code for each agent framework. MCP is an open-source standard for connecting AI applications to external systems, often described as a USB-C port for AI: build an integration once as an MCP server, and any MCP-compatible agent or application can use it.
- MCP host: the AI application (for example, an IDE or an agent). It creates one MCP client per server it connects to.
- MCP server: a program that exposes context and capabilities. Servers offer three primitives: tools (executable functions that perform actions), resources (data that provides context, such as file contents or database records), and prompts (reusable templates).
MCP standardizes the connection. It does not decide how the model uses what it receives, and it is not the same thing as RAG: RAG grounds an answer in retrieved documents, while MCP is the plumbing an agent can use to reach a data source or take an action.
Multi-agent systems and how agents communicate
Complex applications are often split across several specialized agents, each with focused instructions and only the tools its job needs. The patterns differ in who decides the path:
| Pattern | How it works | Who decides the path |
|---|---|---|
| Agents as tools (supervisor) | An orchestrator calls specialist agents as if they were tools, then combines their results | A central orchestrator |
| Swarm | Peer agents hand tasks to each other, sharing one conversation context | The agents themselves |
| Graph | A flowchart of agents with branches and loops | Developer draws it; the model picks a branch at each node |
| Workflow | A fixed set of tasks and dependencies; independent tasks run in parallel | The developer, in advance |
For communication, MCP connects an agent to tools and data, while the Agent-to-Agent (A2A) protocol connects agents that run as separate services, so one team's agent can delegate work to another team's agent. AWS agent services such as AgentCore Runtime and Strands Agents support both.
Multi-agent designs bring focus and parallelism, but not free efficiency: every agent in the chain makes its own model calls, so token costs usually rise, and each agent should still hold only the permissions it needs.
Scenario: A retailer wants one assistant that answers product questions, checks live order status, and processes returns. Returns require updating the order system and emailing a label.
Reflection Question: Which parts of this need tools, which need memory, and would you design it as one agent or as an orchestrator with specialist agents? What would MCP change about how the order system is connected?
💡 Tip: Map the requirement to the building block. Changing or live data → a tool. "Remember me next time" → long-term memory. Connect once, reuse from any agent → MCP. One agent handing work to another team's agent → A2A. Same steps in the same order every time → a workflow.