3.4.3. LLM-as-a-Judge and Business Alignment Metrics
First Principle: Overlap metrics are cheap but shallow, and human review is deep but slow. An LLM judge fills the gap by giving near-human quality judgments at machine scale. But no quality score, however good, proves that an AI application is paying off for the business.
LLM-as-a-Judge
Think of LLM-as-a-judge like a senior editor grading a junior writer's drafts. The editor does not count matching words; they read for correctness, completeness and tone, and they explain each grade.
In an Amazon Bedrock evaluation job that uses a judge model:
- There are two different models: a generator model that answers the prompts and an evaluator model that scores the answers. (You can also bring your own responses from a model outside Bedrock, and the job skips the generation step.)
- You pick built-in metrics or define custom metrics for your business case. The built-in metrics fall into two groups:
- Quality: correctness, completeness, faithfulness (no information beyond the provided context), helpfulness, logical coherence, relevance, following instructions, and professional style and tone.
- Responsible AI: harmfulness, stereotyping, and refusal (whether the response declines to answer).
- The evaluator returns a score and an explanation for every prompt and response pair. The console shows score histograms, and the full report is written to Amazon S3.
Where it fits:
| Approach | Measures | Speed and cost | Main limitation |
|---|---|---|---|
| ROUGE, BLEU, BERTScore | Overlap or similarity with a reference | Fastest, cheapest | Blind to correctness and tone |
| LLM-as-a-judge | Meaning-level quality and safety | Fast, scales to thousands of responses | The judge is itself a model and can be wrong |
| Human evaluation | Anything a person can judge | Slowest, most expensive | Hard to scale |
Because the judge is itself a model, teams commonly spot-check a sample of its scores with human reviewers to make sure it agrees with expert opinion.
Business Objective Alignment Metrics
Technical and judge metrics tell you whether the output is good. Business alignment metrics tell you whether the application is achieving what it was built for:
- Task completion rate: The share of tasks the application finishes successfully, for example, support conversations resolved without escalation to a human agent.
- User satisfaction: How users rate the experience, for example, post-chat ratings or survey scores.
- Cost per interaction: Total inference spend divided by the number of interactions. Because it is normalized, it compares fairly across pilots and workloads of different sizes, and it connects directly to token-based pricing.
A model can score well on faithfulness and coherence and still fail the business: every answer is accurate, yet customers keep asking for a human. Track both kinds of metric together.
Scenario: Two models tie on task completion rate and user satisfaction in a 10,000-conversation pilot. Model A's inference cost is $180 and Model B's is $450. Cost per interaction ($0.018 versus $0.045) makes the decision clear, and it will still be comparable when the next pilot has a different volume.
Reflection Question: Your chatbot's judge-rated helpfulness rose after an update, but its task completion rate fell. What would you investigate first, and which metric should decide whether the update stays?
⚠️ Exam Tip: LLM-as-a-judge = a second model scores and explains each response. Faithfulness, helpfulness and coherence are quality metrics. Task completion rate, user satisfaction and cost per interaction are business metrics. When the question asks whether the application meets business objectives, choose a business metric.