4.2.3. GenAI Gateway Architecture
💡 First Principle: Every enterprise's GenAI workloads eventually converge on the same set of cross-cutting concerns — authentication, rate limiting, cost tracking, model routing, safety enforcement, and observability. A GenAI gateway centralizes these concerns into a single managed layer so individual application teams don't each implement them independently (and inconsistently).
GenAI gateway components:
Cost attribution with resource tagging:
bedrock.create_inference_profile( # once per team: a tagged application inference profile
inferenceProfileName='team-a-claude',
modelSource={'copyFrom': 'arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-3-sonnet-20240229-v1:0'},
tags=[{'key': 'team', 'value': 'team-a'}, {'key': 'cost-center', 'value': 'CC-1042'}]
)
team = event['requestContext']['authorizer']['team']
bedrock_runtime.invoke_model(modelId=profile_arn_for(team), body=json.dumps(payload)) # InvokeModel has no tags parameter
Activate the tag as a cost allocation tag to split Bedrock spend by team in Cost Explorer. For near-real-time warnings, have the gateway publish each team's token counts (from the response usage field) as a CloudWatch custom metric with a team dimension and alarm at, for example, 80% of that team's budget. This alerts without blocking; the authorizer's rate limit below is the blocking control.
Rate limiting per team — Lambda authorizer pattern:
def lambda_authorizer(event, context):
api_key = event['headers'].get('x-api-key')
team = lookup_team_from_key(api_key)
# Check rate limit in ElastiCache (token bucket per team)
remaining_tokens = check_and_decrement_rate_limit(team)
if remaining_tokens <= 0:
raise Exception("Unauthorized") # Returns 401 to client
return generate_allow_policy(team, event['methodArn'])
Multi-tenant SaaS on the gateway: give each customer an API key tied to an API Gateway usage plan (per-customer throttling and quotas), and let the Lambda authorizer resolve the tenant and pass it in the request context. Store per-tenant configuration (system prompt, knowledge base ID, guardrail ID and version) in DynamoDB keyed by tenant ID and look it up per request, rather than running separate accounts or infrastructure per customer.
⚠️ Exam Trap: Building a GenAI gateway on API Gateway + Lambda has a 29-second timeout ceiling. For streaming FM responses that exceed this, the gateway layer must use API Gateway WebSocket APIs or a streaming-capable proxy (Lambda Function URLs with response streaming) rather than the standard REST API integration.
Reflection Question: Your organization has 8 application teams each making direct Bedrock API calls. In a quarterly review, you discover three teams have no content filtering, two teams are using the most expensive model for simple classification tasks, and cost attribution is impossible because all calls share one IAM role. Design the minimal GenAI gateway that fixes all three problems.