30% off every course until Sunday, October 11. Our biggest update yet, and we'd like you to try it. Applied automatically at checkout.

Choose your certification
Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

2.1.5. Context Engineering and Token-Based Pricing

First Principle: A foundation model can only work with what is in its context window for that request, and you pay (in money and in time) for every token that goes in and comes out. Deciding what to put in the context and how much the model should produce is therefore both a quality decision and a cost decision.

Context Engineering

AWS defines context engineering as the practice of designing systems that dynamically assemble the optimal set of information for an LLM to perform a given task. Where prompt engineering focuses on what to ask, context engineering focuses on what to show: the context window becomes a workspace that is filled with the right components for each request.

Typical components of an engineered context:

  • System prompt or instructions: the model's role, the task and the output format, often with one-shot or few-shot examples.
  • User query: the actual request.
  • User profile: facts about this user, such as their subscription level or history.
  • Memory: relevant earlier turns of the conversation.
  • Tool definitions and Model Context Protocol (MCP) servers: ways for the model to look up live data, such as a customer record in a CRM.
  • Knowledge bases: documents retrieved with RAG to ground the answer in facts.

Teams usually discover the need for context engineering the hard way: clear instructions still produce generic answers because the model was never shown the customer's data or the relevant documents. Better wording cannot fill that gap; better context can. Context engineering does not change the model's weights, and inference parameters such as temperature are settings, not context.

Because every assembled component costs input tokens, good context engineering also means leaving out what the task does not need: the relevant passages rather than the whole document set, the relevant turns rather than the entire history.

Token-Based Pricing

With on-demand inference (for example, on Amazon Bedrock), you are billed per token, with separate prices for input tokens and output tokens that vary by model and Region. Output tokens are often priced several times higher than input tokens, so a short prompt that asks for a long report spends most of its cost on the output.

Ways the pricing model shapes design:

  • Ask for only the output you need and set a sensible maximum output length (max_tokens).
  • Keep repeated static content cacheable. Prompt caching bills tokens read from the cache at a reduced rate on supported models (on some models, writing tokens to the cache costs somewhat more than a normal input token, so caching pays off when the prefix is reused).
  • Move non-urgent bulk work to batch inference, which is priced lower than on-demand for select models (see 2.3.2).
Tokens and Performance

Token counts also drive latency. A request passes through two stages:

  • Prefill: the model processes the whole input prompt before producing the first token, so input length mainly drives the time to first token.
  • Decode: output tokens are generated one at a time, so total response time grows with the number of output tokens.

Tokens also count against your account's tokens-per-minute quota. On Amazon Bedrock, the request's input tokens plus its max_tokens value are reserved from the quota when the request starts, so setting max_tokens far above what you need reduces how many requests can run at once. You are billed only for the tokens actually used.

SymptomLikely token causeTypical fix
High cost, short prompts, long answersOutput tokensAsk for concise output; cap max_tokens
Slow to start respondingVery long inputTrim or cache the static part of the prompt
Slow to finish respondingLong outputShorter answers; cap max_tokens
Requests throttled earlier than expectedmax_tokens set far too highRight-size max_tokens to the expected output

Scenario: A support assistant sends the full 200-page product manual with every question "so the model has everything", and asks for detailed answers. It is expensive, slow, and still misses details about each customer's own order.

Reflection Question: How would context engineering change what is sent (retrieved passages, the customer's order record, relevant history), and how would that change both answer quality and token cost?

💡 Tip: Prompt engineering = what to ask. Context engineering = what to show. Token-based pricing = what it costs, split into input and output. Long answers cost the most and take the longest.

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications