2.1.5. Context Engineering and Token-Based Pricing
First Principle: A foundation model can only work with what is in its context window for that request, and you pay (in money and in time) for every token that goes in and comes out. Deciding what to put in the context and how much the model should produce is therefore both a quality decision and a cost decision.
Context Engineering
AWS defines context engineering as the practice of designing systems that dynamically assemble the optimal set of information for an LLM to perform a given task. Where prompt engineering focuses on what to ask, context engineering focuses on what to show: the context window becomes a workspace that is filled with the right components for each request.
Typical components of an engineered context:
- System prompt or instructions: the model's role, the task and the output format, often with one-shot or few-shot examples.
- User query: the actual request.
- User profile: facts about this user, such as their subscription level or history.
- Memory: relevant earlier turns of the conversation.
- Tool definitions and Model Context Protocol (MCP) servers: ways for the model to look up live data, such as a customer record in a CRM.
- Knowledge bases: documents retrieved with RAG to ground the answer in facts.
Teams usually discover the need for context engineering the hard way: clear instructions still produce generic answers because the model was never shown the customer's data or the relevant documents. Better wording cannot fill that gap; better context can. Context engineering does not change the model's weights, and inference parameters such as temperature are settings, not context.
Because every assembled component costs input tokens, good context engineering also means leaving out what the task does not need: the relevant passages rather than the whole document set, the relevant turns rather than the entire history.
Token-Based Pricing
With on-demand inference (for example, on Amazon Bedrock), you are billed per token, with separate prices for input tokens and output tokens that vary by model and Region. Output tokens are often priced several times higher than input tokens, so a short prompt that asks for a long report spends most of its cost on the output.
Ways the pricing model shapes design:
- Ask for only the output you need and set a sensible maximum output length (
max_tokens). - Keep repeated static content cacheable. Prompt caching bills tokens read from the cache at a reduced rate on supported models (on some models, writing tokens to the cache costs somewhat more than a normal input token, so caching pays off when the prefix is reused).
- Move non-urgent bulk work to batch inference, which is priced lower than on-demand for select models (see 2.3.2).
Tokens and Performance
Token counts also drive latency. A request passes through two stages:
- Prefill: the model processes the whole input prompt before producing the first token, so input length mainly drives the time to first token.
- Decode: output tokens are generated one at a time, so total response time grows with the number of output tokens.
Tokens also count against your account's tokens-per-minute quota. On Amazon Bedrock, the request's input tokens plus its max_tokens value are reserved from the quota when the request starts, so setting max_tokens far above what you need reduces how many requests can run at once. You are billed only for the tokens actually used.
| Symptom | Likely token cause | Typical fix |
|---|---|---|
| High cost, short prompts, long answers | Output tokens | Ask for concise output; cap max_tokens |
| Slow to start responding | Very long input | Trim or cache the static part of the prompt |
| Slow to finish responding | Long output | Shorter answers; cap max_tokens |
| Requests throttled earlier than expected | max_tokens set far too high | Right-size max_tokens to the expected output |
Scenario: A support assistant sends the full 200-page product manual with every question "so the model has everything", and asks for detailed answers. It is expensive, slow, and still misses details about each customer's own order.
Reflection Question: How would context engineering change what is sent (retrieved passages, the customer's order record, relevant history), and how would that change both answer quality and token cost?
💡 Tip: Prompt engineering = what to ask. Context engineering = what to show. Token-based pricing = what it costs, split into input and output. Long answers cost the most and take the longest.