An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

5.1.1. Amazon Bedrock Guardrails

💡 First Principle: Bedrock Guardrails operates as a managed content safety layer that sits between your application and the FM — it evaluates every input before it reaches the model and every output before it reaches the user, enforcing your defined policies without requiring custom code.

Guardrails configuration — the six protection categories:
CategoryWhat It DoesConfiguration
Topic denialBlock the FM from discussing specific topicsNatural language topic descriptions
Content filtersFilter harmful content (hate, insults, sexual, violence, misconduct) plus prompt attacks (jailbreak, injection)Strength per category (NONE/LOW/MEDIUM/HIGH)
Word filtersBlock specific words, phrases, or regex patternsCustom word lists + profanity filter
PII redactionDetect and mask/block personal dataChoose PII entity types; action = BLOCK or ANONYMIZE
GroundingVerify FM response is factually supported by retrieved contextThreshold score (0–1); block responses below threshold
Sensitive informationCustom regex patterns for domain-specific sensitive dataRegex patterns + action
Applying Guardrails to a Bedrock invocation:
response = bedrock_runtime.converse(
    modelId='anthropic.claude-3-sonnet-20240229-v1:0',
    guardrailConfig={
        'guardrailIdentifier': 'arn:aws:bedrock:us-east-1:123456789:guardrail/GUARDRAILID',
        'guardrailVersion': 'DRAFT',  # Or specific version number
        'trace': 'ENABLED'  # Returns trace showing which policy triggered
    },
    messages=[{'role': 'user', 'content': [{'text': user_input}]}]
)

# Check if Guardrails blocked the response
if response['stopReason'] == 'guardrail_intervened':
    guardrail_trace = response['trace']['guardrail']
    triggered_policy = guardrail_trace['inputAssessment']['topicPolicy']['topics'][0]
    log_security_event(triggered_policy, user_input)
    return "I'm not able to help with that topic."

Standalone checks with ApplyGuardrail: the ApplyGuardrail API evaluates text against a guardrail without invoking a Bedrock model. Use it to apply one policy to output from self-hosted, SageMaker, or third-party models, or at any pipeline stage such as retrieved chunks; a guardrail attached to a model call covers only that Bedrock invocation.

Defense-in-depth with Guardrails + Comprehend + Lambda:

⚠️ Exam Trap: Guardrails with trace: ENABLED returns detailed information about which policy triggered and why — but this trace data includes the blocked content. Logging trace data to CloudWatch Logs creates a record of harmful content that users attempted to input. Your log retention and access control policies must account for this security-sensitive data in your logs.

Tuning filter strength: strength sets how confident the classifier must be before content is blocked. LOW blocks only high-confidence harm; MEDIUM blocks high and medium; HIGH also blocks low-confidence harm. Borderline content can therefore pass a MEDIUM filter, while HIGH over-blocks legitimate material (a history lesson tripping VIOLENCE). HIGH is the strictest setting. Tune strength per category. To keep specific subjects blocked after relaxing a filter, add narrow denied topics. Denied topics only block; there is no allow-list that exempts content from a filter. If harmful content still slips past HIGH, add an independent output check, such as Amazon Comprehend toxicity detection in a post-processing Lambda.

Guardrails are per-request resources: each guardrail (with its versions) is a separate resource, selected per call through guardrailIdentifier. Different user tiers or apps can use different guardrails from the same code. When a guardrail intervenes, the caller receives the blocked-input or blocked-output message you configured instead of the content.

Custom regex patterns match shape, not meaning: a sensitive-information pattern for account numbers that accepts any run of digits also matches the digits in a balance such as "$45,231.00", so the balance is masked too. Anchor custom patterns tightly (length, prefix, word boundaries) and test them against legitimate responses.

Reflection Question: A competitor analysis chatbot should never discuss your company's revenue figures or employee headcount (confidential). Users have discovered they can extract this information by asking the FM to "roleplay as a financial analyst" or "pretend you're writing a fictional story about a company like ours." What Guardrails configuration addresses this, and why does topic denial handle indirect attacks better than word filters?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications