2.1.2. The Bedrock Model Catalog: Claude, Titan, Llama, and Beyond
💡 First Principle: The Bedrock model catalog is organized by provider and capability tier — understanding which providers offer which capabilities, and which models within each provider suit which task types, is prerequisite to any selection decision.
Primary model families on Amazon Bedrock (AIP-C01 relevant):
| Provider | Model Family | Strengths | Context Window | Best For |
|---|---|---|---|---|
| Anthropic | Claude 3 Haiku | Fast, low cost | 200K tokens | High-volume simple tasks, classification |
| Anthropic | Claude 3 Sonnet | Balanced | 200K tokens | General-purpose reasoning, code |
| Anthropic | Claude 3 Opus | Highest capability | 200K tokens | Complex reasoning, long docs |
| Amazon | Titan Text Lite | Cost-optimized | 4K tokens | Simple tasks, tight budgets |
| Amazon | Titan Text Premier | Balanced | 32K tokens | General enterprise use |
| Amazon | Titan Embeddings v2 | 1024-dim embeddings | — | RAG, semantic search |
| Amazon | Titan Multimodal | Image + text | — | Product catalog, image Q&A |
| Meta | Llama 3 (various) | Open weights, customizable | 8K–128K | When open-source licensing needed |
| Mistral AI | Mistral/Mixtral | Efficient, multilingual | 32K tokens | European data residency requirements |
| Stability AI | Stable Diffusion XL | Image generation | — | Creative content, product images |
Amazon Bedrock Marketplace: 100+ additional specialized and emerging models (e.g., domain-specific medical or financial models) beyond the serverless catalog. You subscribe, then deploy the model to a SageMaker AI–managed endpoint on AWS infrastructure, choosing the instance type and count; you pay the provider's software fee plus the endpoint's infrastructure cost. Deployed models are called through the Bedrock InvokeModel/Converse APIs and work with Agents, Knowledge Bases and Guardrails.
Cross-Region Inference: When a model is not available in your required AWS region (common for new model releases), Bedrock's cross-region inference automatically routes your request to the nearest region where the model is available. This is transparent to your application — you use the same API call with a cross-region inference profile ARN.
# Cross-region inference profile — handles routing automatically
response = bedrock_runtime.invoke_model(
modelId='us.anthropic.claude-3-5-sonnet-20241022-v2:0', # US inference profile ID, not a foundation-model ARN
# Bedrock may serve the request from any US Region in the profile (e.g., us-east-1, us-east-2, us-west-2), never outside the US
body=json.dumps({'messages': [...], 'max_tokens': 1000})
)
⚠️ Exam Trap: Cross-region inference means your data may leave your primary AWS region to be processed in another region. For workloads with strict data residency requirements (GDPR, HIPAA data that must stay in eu-west-1), cross-region inference must be disabled or configured with region constraints — for EU-only residency, a geographic inference profile (the eu. prefix) routes requests only among EU Regions, while a US or global profile would move the data out of the EU. The exam specifically tests this data residency conflict.
Reflection Question: A European healthcare company processes patient data and wants to use a foundation model available only in us-east-1. They have strict GDPR data residency requirements preventing patient data from leaving the EU. What is the correct architectural approach?