3.1.2. Amazon Bedrock Knowledge Bases Architecture
💡 First Principle: Bedrock Knowledge Bases operationalizes the full RAG pipeline as a managed service — it moves the undifferentiated heavy lifting of ingestion, chunking, embedding, indexing, and retrieval out of your application code and into a fully managed AWS service. You define what to index; AWS manages how.
The Bedrock Knowledge Bases architecture:
Sync configuration — the most frequently tested operational detail: Knowledge Bases does not automatically update when source documents change. You must trigger a sync:
- Manual sync: Via console or API call (
bedrock-agent.start_ingestion_job()) - Scheduled sync: EventBridge Scheduler triggers sync job on a cron schedule
- Event-driven sync: S3 event notification → Lambda → start ingestion job on document change
- Troubleshooting: each ingestion job reports scanned/indexed/failed document counts and per-document failure reasons (unsupported format, file too large, etc.) — read those before changing the configuration
- API error codes:
AccessDeniedException(403) means the caller lacks permission (e.g.,bedrock:StartIngestionJob);ResourceNotFoundException(404) means the Knowledge Base ID was not found — usually a wrong ID or a client pointed at a different Region, since Knowledge Bases are Regional
# Event-driven sync triggered by S3 object creation
def lambda_handler(event, context):
bedrock_agent = boto3.client('bedrock-agent')
response = bedrock_agent.start_ingestion_job(
knowledgeBaseId='KBID123456',
dataSourceId='DSID789012',
description=f"Auto-sync triggered by {event['Records'][0]['s3']['object']['key']}"
)
return response['ingestionJob']['ingestionJobId']
Metadata schema for filtered retrieval:
Documents in S3 can have accompanying .metadata.json files that define structured attributes for filtered retrieval:
{
"metadataAttributes": {
"department": "legal",
"document_type": "policy",
"effective_date": "2024-01-01",
"confidentiality": "internal"
}
}
This enables queries like "retrieve only documents from the legal department effective after 2024" — combining semantic similarity with structured filtering.
Filters only match chunks that were indexed with the attribute. If you add a new metadata field later, documents ingested before it won't match filters on that field until their .metadata.json is updated and the data source is re-synced. With OpenSearch Serverless you don't pre-declare the attribute. With Aurora, filtering needs the JSONB custom_metadata column (or a column per attribute), set up before the knowledge base is created.
⚠️ Exam Trap: Metadata filters narrow relevance; they are not a security boundary — the application must attach the right filter to every query, and one missed filter leaks another tenant's content. When tenants must never see each other's documents, give each tenant its own Knowledge Base and vector index (silo model).
⚠️ Exam Trap: Bedrock Knowledge Bases sync jobs are not instantaneous — each one scans the data source and can take minutes to hours for large corpora. When documents must be searchable within seconds of upload, don't wait for a sync: ingest each document as it arrives, either through the Knowledge Base's direct-ingestion API (IngestKnowledgeBaseDocuments, for S3 and custom data sources; for S3, also write the file to the bucket so the next sync doesn't undo it) or with a custom pipeline (S3 event → Lambda → embed → OpenSearch).
Querying a knowledge base — two runtime APIs:
RetrieveAndGenerateruns retrieval and generation in one call, which makes it the least-code way to put a knowledge base behind a bot or API. Its response carries acitationsarray (each generated span linked to itsretrievedReferences) and asessionId. Pass thesessionIdback on the next call and Bedrock keeps the conversation context for multi-turn follow-ups. You can also customize its generation prompt template, for example to control tone or how sources are mentioned.Retrievereturns only the ranked chunks with their source metadata. Use it when custom logic must sit between retrieval and generation, such as filtering chunks by the caller's clearance, redacting, or reranking in your own code (e.g., Lambda). Then call the FM yourself with Converse.
Changing the vector store (e.g., Aurora pgvector → OpenSearch Serverless) is not a migration Bedrock performs. Create a knowledge base on the new store and run a full ingestion job from the unchanged S3 sources, which re-chunks, re-embeds and re-indexes every document.
Reflection Question: A compliance team uploads a new regulatory document to S3 and expects the chatbot to be able to answer questions about it "immediately." They're currently using Bedrock Knowledge Bases with a nightly sync job. How would you re-architect the pipeline to minimize the delay between document upload and query availability?