Skip to content

Prompt Evaluation

Your prompt produces great results 80% of the time. But which 20% fails? And why?

Without evaluation, you’re flying blind. You can’t:

  • Know if a prompt change improved quality
  • Compare different prompt strategies objectively
  • Detect regressions before users do
  • Prove quality to stakeholders

Prompt evaluation exists because:

  • Prompts are probabilistic — same prompt, different results
  • Quality is subjective — but needs objective measurement
  • Regressions happen — changes break existing behavior
  • ROI needs proof — justify prompt engineering effort
  • Production needs guardrails — enforce quality standards

“If you can’t measure it, you can’t improve it.” — Peter Drucker (applied to prompts)


Scenario: Your team builds a product search prompt.

Version A: “Return products matching the query.”

  • Works for “red shoes”
  • Fails for “birthday gift under $50 for mom”

Version B: “Recommend products for the user’s intent and budget.”

  • Works for both
  • But sometimes returns irrelevant results

Without evaluation: You guess which is better. With evaluation: You measure both on 100 test cases and know objectively.


MetricWhat It MeasuresHow to Measure
CorrectnessFactually accurateCompare to ground truth
RelevanceAnswers the questionHuman rating / LLM score
CompletenessCovers all requirementsChecklist
ConcisenessNo unnecessary infoToken count
GroundednessBased on provided contextCheck against source
SafetyNo harmful contentModeration API
ConsistencySame input → similar outputRun N times, measure variance
MetricWhat It Measures
Task Success RateDid the user complete their goal?
User SatisfactionRatings, feedback, retention
Cost per QueryToken usage × model price
LatencyTime from request to response
Escalation RateHow often human intervention needed

flowchart TD
subgraph Inputs
A[Prompt Version]
B[Test Dataset]
C[Ground Truth]
end
subgraph Evaluation
D[Run Inference]
E[Metrics Calculation]
F{Auto or Human?}
end
subgraph Auto
G[LLM-as-a-Judge]
H[Checklist Checker]
I[Regex/Validation]
end
subgraph Human
J[Raters]
K[Experts]
L[User Feedback]
end
subgraph Output
M[Score Report]
N[Regression Alert]
O[Improvement Suggestions]
end
A --> D
B --> D
D --> E
C --> E
E --> F
F -->|Auto| G
F -->|Auto| H
F -->|Auto| I
F -->|Human| J
F -->|Human| K
F -->|Human| L
G --> M
H --> M
I --> M
J --> M
K --> M
L --> M
M --> N
M --> O

Use an LLM to evaluate prompt outputs.

evaluation_prompt = """
You are an expert evaluator. Rate the following response on:
- Accuracy (1-5): Is the information correct?
- Relevance (1-5): Does it answer the query?
- Completeness (1-5): Does it cover everything?
Query: {query}
Response: {response}
Return JSON only:
{"accuracy": 4, "relevance": 5, "completeness": 3}
"""
checklist:
response_includes_answer: true
response_mentions_price: true
response_under_200_tokens: true
response_no_hallucinations: true
response_formatted_correctly: true
human_eval:
method: "pairwise_comparison"
raters: 3
scale: "A > B, B > A, or Equal"
criteria:
- helpfulness
- accuracy
- tone

sequenceDiagram
participant User
participant App
participant LLM
participant Judge
participant Report
User->>App: Query
App->>LLM: Send prompt
LLM-->>App: Response
App->>User: Response
par Evaluation
App->>Judge: Send query + response
Judge->>Judge: Score on metrics
Judge-->>Report: Accuracy: 4, Relevance: 5
Report->>Report: Log to database
end
Note over Judge,Report: Runs async in background

[
{
"id": "test-001",
"query": "What is the return policy?",
"expected": "30-day return window, unopened items",
"context": "E-commerce support",
"difficulty": "easy",
"category": "policy"
},
{
"id": "test-002",
"query": "My order hasn't arrived in 2 weeks. What should I do?",
"expected": "Apologize, offer tracking lookup, explain escalation path",
"context": "E-commerce support",
"difficulty": "medium",
"category": "issue"
}
]
Use CaseMinimum SizeRecommended
Development10-20Quick iteration
Regression Testing50-100Core scenarios
Production Release200-500Comprehensive
Model Comparison500-1000Statistical significance

Bad PracticeGood Practice
”Feels better""Score improved 12%“
Evaluate onceRegular automated evaluation
Single metricBalanced scorecard
Same test alwaysRotating + regression tests
No golden datasetVersioned golden dataset

flowchart LR
A[Push Prompt Change] --> B[Unit Tests]
B --> C[Eval Suite]
C --> D{Quality Gate}
D -->|Pass| E[Staging]
D -->|Fail| F[Blocked + Notification]
E --> G[Canary]
G --> H{Monitor}
H -->|OK| I[Production]
H -->|Issue| J[Rollback]
style D fill:#eab308,color:#000
style F fill:#ef4444,color:#fff
style I fill:#22c55e,color:#000
style J fill:#ef4444,color:#fff
quality_gates:
accuracy:
minimum: 0.85
action: block
relevance:
minimum: 0.80
action: warn
safety:
minimum: 0.99
action: block
latency:
maximum: "2000ms"
action: warn

ToolTypeBest For
LangSmithManagedEvaluation + tracing
Weights & BiasesManagedExperiment tracking
PromptLayerManagedLogging + evaluation
DeepEvalOpen SourceUnit testing for LLMs
RagasOpen SourceRAG evaluation
EleutherAI LM EvalOpen SourceModel evaluation

MistakeWhy It HurtsFix
No evaluation at allBlind changesStart with 10 test cases
Bias in evaluationJudge favors certain stylesUse multiple judges
Overfitting to datasetDoesn’t generalizeRotate test cases
Only auto-evaluationMiss nuanceAdd human sampling
Ignoring edge casesUnexpected failuresInclude adversarial tests

PracticeDescription
Version datasetsGolden dataset changes over time
Multiple judgesUse 3+ LLM judges, take majority
Sample human eval10% of production calls manually reviewed
Track over timeMetrics dashboard with history
Calibrate judgesCheck LLM judge against human ratings
Test regressions firstRun existing tests before new ones

  1. What is prompt evaluation and why is it important?
  2. Name three metrics to evaluate prompt quality.
  1. How does LLM-as-a-Judge work? What are its limitations?
  2. Design an evaluation dataset for a customer support chatbot.
  1. How would you detect prompt regressions automatically?
  2. Design a quality gate system for prompt deployments.
  1. How do you ensure evaluation is not biased toward specific prompt styles?
  2. Design a company-wide prompt evaluation framework across multiple teams and use cases.

  • Measure what matters — choose metrics aligned with business goals
  • Automate evaluation — catch regressions before users do
  • Combine methods — auto for scale, human for nuance
  • Maintain datasets — golden datasets evolve with product
  • Gate deployments — enforce quality thresholds

Key Insight: Good evaluation turns prompt engineering from an art into an engineering discipline.


Next: Document 22 — Prompt Security