5.1.4. Hallucination Detection and Grounding
First Principle: A model generates text that is statistically likely, not text that is verified. Improving output accuracy therefore takes two layers: ground the model in trusted information before it answers, then check the answer before anyone relies on it.
Think of it like a newsroom. The reporter is handed the source files (grounding), and an editor checks the story against those files before it runs (validation). Neither step alone is enough.
Grounding techniques (prevention)
- RAG grounding: Retrieve relevant passages from a trusted knowledge base (for example, Amazon Bedrock Knowledge Bases) and supply them with the prompt, so the model answers from your documents rather than from what it may misremember.
- Source citation: Return the passages an answer was based on, so users and reviewers can verify it (see 5.1.3).
- Lower temperature for factual tasks: less randomness means less invention.
Detection techniques (validation)
- Output validation: Check each response against the source it should be based on, or against rules it must follow, before it is shown.
- Confidence scoring: Score how well supported a response is, and block or flag anything below a threshold.
AWS implementation: Amazon Bedrock Guardrails
-
Contextual grounding checks need three things: a grounding source, the user's query and the model's response. They score two paradigms:
- Grounding: Is the response factually supported by the source? Any new information is treated as ungrounded.
- Relevance: Does the response answer the user's query?
Each response receives confidence scores, and you set thresholds between 0 and 0.99. Responses scoring below a threshold are detected as hallucinations and blocked. Raising the threshold blocks more ungrounded content but can also filter some acceptable answers; a threshold of 1 is invalid because it would block everything. Supported use cases are summarization, paraphrasing and question answering; AWS lists conversational QA and chatbot use cases as not supported.
-
Automated Reasoning checks validate responses against a set of logical rules (for example, an encoded policy). They detect hallucinations, suggest corrections and highlight unstated assumptions.
Example: the grounding source says "London is the capital of the UK. Tokyo is the capital of Japan." The user asks, "What is the capital of Japan?"
| Model response | Grounding | Relevance |
|---|---|---|
| "The capital of Japan is Tokyo." | High | High |
| "The capital of Japan is London." | Low | High |
| "The capital of the UK is London." | High | Low |
Scenario: A RAG application summarizes internal policy documents, and reviewers keep finding figures that are not in the retrieved text. Adding a guardrail with a contextual grounding check, using the retrieved passages as the grounding source, blocks those summaries before users see them.
Reflection Question: Your team wants to catch HR answers that contradict the company's leave-eligibility rules. Would you use contextual grounding checks or Automated Reasoning checks, and what would you need to prepare for each?
⚠️ Exam Tip: RAG reduces hallucinations; it does not eliminate them. Pair grounding with validation. Source document + query → contextual grounding checks. Logical rules → Automated Reasoning checks. Grounded but off-topic → low relevance, not low grounding.