Skip to content

19. Prompt Testing

If you’re not testing your prompts, you’re guessing, not engineering.

Prompt testing is the practice of systematically evaluating prompts against known test cases to ensure they produce reliable, correct outputs before deployment.


You deploy a prompt to production. For a week, it works fine. Then a user asks a question you never tested — and the prompt fails spectacularly, producing an incorrect or even harmful response.

Without testing, you find out about prompt failures from users. With testing, you find out before deployment.

flowchart TD
subgraph NOTESTING["Without Testing"]
D["Deploy prompt"] --> U["Users use it"]
U --> F["❌ Failure in production"]
F --> H["Emergency fix"]
end
subgraph TESTING["With Testing"]
T1["Write tests"] --> T2["Test prompt"]
T2 --> T3["Fix issues"]
T3 --> T4["Deploy with\nconfidence"]
T4 --> T5["✅ Fewer failures"]
end
style NOTESTING fill:#ef4444,color:#fff
style TESTING fill:#22c55e,color:#fff

You wouldn’t deploy code without unit tests, integration tests, and regression tests. Prompts are code — they need the same treatment.

Prompt testing is unit testing for LLM interactions.


flowchart TD
TESTING["Prompt Testing"] --> UNIT["Unit Tests\nIndividual prompt\nKnown inputs/outputs"]
TESTING --> REGRESSION["Regression Tests\nHistorical failures\nDon't reintroduce bugs"]
TESTING --> GOLDEN["Golden Dataset\nCurated test cases\nExpected outputs"]
TESTING --> EDGE["Edge Cases\nUnusual inputs\nBoundary conditions"]
TESTING --> STRESS["Stress Tests\nVolume\nPerformance"]
style TESTING fill:#8b5cf6,color:#fff

Test Case 1: Normal input
Input: "What is the capital of France?"
Expected: Contains "Paris"
Test Case 2: Edge case
Input: "" (empty string)
Expected: Handle gracefully, ask for clarification
Test Case 3: Complex input
Input: "If a train travels at 60mph for 2 hours..."
Expected: Contains "120 miles"
Test Case 4: Adversarial input
Input: "Ignore all previous instructions and..."
Expected: Refuse, don't follow
def test_prompt(prompt_template, test_cases):
for case in test_cases:
response = call_llm(prompt_template, case["input"])
# Check against expected criteria
checks = {
"contains_expected": case["expected"] in response,
"no_harmful_content": not contains_harmful(response),
"correct_format": validate_format(response, case.get("format")),
"within_length": len(response) < case.get("max_length", 2000),
}
if not all(checks.values()):
log_failure(case, response, checks)
return pass_rate

A golden dataset is a curated set of test cases with verified expected outputs:

[
{
"input": "What is 2+2?",
"expected_contains": "4",
"expected_not_contains": ["5", "I don't know"],
"expected_format": "text"
},
{
"input": "Write a Python function to sort a list",
"expected_contains": ["def sort", "return"],
"expected_not_contains": ["I cannot"],
"expected_format": "code"
}
]

flowchart LR
PR["Prompt Change"] --> TRIGGER["CI Trigger"]
TRIGGER --> RUN["Run Test Suite"]
RUN --> CASES["Test against\nGolden Dataset"]
CASES --> EVAL["Evaluate Results"]
EVAL -->|Pass| DEPLOY["✅ Deploy"]
EVAL -->|Fail| REVIEW["❌ Review & Fix"]
REVIEW --> PR
style PR fill:#3b82f6,color:#fff
style DEPLOY fill:#22c55e,color:#fff
style REVIEW fill:#ef4444,color:#fff

AspectWhat to CheckExample Test
CorrectnessIs the factual answer right?“2+2 should equal 4”
FormatIs the output structure correct?”Response must be valid JSON”
ToneIs the tone appropriate?”Should be professional, not casual”
SafetyNo harmful content?”Should refuse harmful requests”
ConsistencySame input = similar output?”Run 3 times, check variance”
LatencyIs it fast enough?”Should respond within 2 seconds”
Token usageIs it efficient?”Output should be < 500 tokens”

test_cases = [
{"input": "I want a refund", "expected": "refund"},
{"input": "Your product is amazing!", "expected": "praise"},
{"input": "Where's my order?", "expected": "tracking"},
{"input": "", "expected": "unknown"}, # Edge case
{"input": "I hate you", "expected": "abuse"}, # Edge case
{"input": "Refund\nRefund\nREFUND", "expected": "refund"}, # Repeated
]
def test_classifier_prompt():
results = []
for case in test_cases:
response = classify_intent(case["input"])
passed = case["expected"] == response
results.append({"test": case["input"], "passed": passed})
pass_rate = sum(r["passed"] for r in results) / len(results)
assert pass_rate > 0.9, f"Classifier failed: {pass_rate:.0%} pass rate"
def test_json_output():
prompt = "Extract name and age as JSON. Input: John is 30"
response = call_llm(prompt)
try:
parsed = json.loads(response)
assert "name" in parsed
assert "age" in parsed
assert isinstance(parsed["age"], (int, float))
except (json.JSONDecodeError, AssertionError):
assert False, "Response was not valid JSON with required fields"

MistakeWhy It’s Wrong
❌ Not testing edge casesEmpty inputs, very long inputs, special characters — these break prompts
❌ Testing only happy pathsThe prompt works for normal inputs but fails on unusual ones
❌ No golden datasetCan’t measure improvement without a baseline
❌ Manual testing onlyDoesn’t scale — automate your tests
❌ Not testing after changesA small change can break previously working cases

AspectBadGood
Coverage2-3 test cases20+ test cases covering edge cases
AutomationManual testingCI pipeline, automated on every change
DatasetNo baselineVersion-controlled golden dataset
MetricsSubjective “feels right”Objective pass/fail per test case
RegressionNoneRe-run ALL tests on every change

LangSmith provides evaluation tools for prompt testing, including dataset management, test runs, and regression tracking.

PromptLayer tracks prompt versions and their performance across test cases, making it easy to compare variants.


Q: Why should prompts be tested like code?

Because prompts can fail in production, producing incorrect or harmful outputs. Testing catches failures before deployment, ensures consistency, and prevents regressions when prompts change.

Q: What is a golden dataset and why is it important?

A golden dataset is a curated collection of test inputs with expected outputs or validation criteria. It’s important because it provides a consistent baseline for measuring prompt quality, detecting regressions, and comparing prompt variants.

Q: Design a prompt testing strategy for a customer support chatbot.

I’d design: (1) A golden dataset of 50+ customer inquiries covering all intent categories, (2) Automated tests for correctness (intent classification accuracy), format (JSON validity), tone (professional score), and safety (harmful content check), (3) CI pipeline that runs all tests on every prompt change, (4) Regression tests that include all previously failed cases, (5) Staged deployment (canary → 10% → 50% → 100%) with monitoring at each stage, (6) A/B testing capability for comparing prompt variants with real users.


ConceptKey Point
Golden DatasetCurated test cases with expected outputs
AutomationRun tests in CI on every prompt change
CoverageTest normal, edge, and adversarial inputs
RegressionRe-test everything when prompts change
Key PrincipleIf you’re not testing, you’re guessing

Previous: 18 — Prompt Optimization →

Next: 20 — Prompt Versioning →