Software Engineering Without Unit Tests is Chaos
In traditional software development, no competent team would deploy code to production without continuous integration: unit tests, integration tests, and linting checks that must pass before merging a pull request.
Yet in generative AI, engineering teams routinely modify system prompts, upgrade model checkpoints (e.g. from gpt-4o-2024-05-13 to gpt-4o-2024-08-06), or tweak temperature values based on nothing more than 2 manual checks in a playground UI.
Two days later, production error rates spike because the new prompt broke structured JSON outputs for 12% of mobile users.
Evals-Driven Development (EDD) brings software engineering discipline to AI: you define objective assertion suites and run automated regression benchmarks in GitHub Actions before any prompt change merges into production.
1. The Anatomy of an LLM Eval
An evaluation suite tests non-deterministic systems using three distinct layers of assertions:
graph TD
Output[Model Generation] --> L1[Layer 1: Deterministic Schema & Type Assertions]
Output --> L2[Layer 2: Rule-Based Heuristic Tests (Regex, Length, Blacklist)]
Output --> L3[Layer 3: Model-as-a-Judge Semantic Scoring (Faithfulness, Coherence)]
L1 --> Pass[Pass / Fail Gate in CI]
L2 --> Pass
L3 --> Pass
- Layer 1 (Deterministic): Did the output conform strictly to the Zod schema? Is it valid JSON? Does it contain required keys?
- Layer 2 (Heuristic): Did the output avoid blacklisted toxic terms? Is the token length within budget bounds? Did it include mandatory disclaimer strings?
- Layer 3 (Semantic LLM-as-a-Judge): Given the reference documentation, is the answer strictly faithful (free of hallucinations)?
2. Implementing Automated Evals with GitHub Actions
Here is a practical eval runner using TypeScript and Vitest that runs in CI on every pull request:
// evals/prompt-regression.test.ts
import { describe, it, expect } from 'vitest';
import { generateCustomerSummary } from '../src/ai/summarizer';
import goldenDataset from './golden-dataset.json';
describe('Customer Support Summarizer Regression Suite', () => {
goldenDataset.forEach((fixture) => {
it(`evaluates test case: ${fixture.id} - ${fixture.scenario}`, async () => {
const result = await generateCustomerSummary(fixture.transcript);
// 1. Deterministic checks
expect(result).toHaveProperty('sentiment');
expect(['positive', 'neutral', 'negative']).toContain(result.sentiment);
expect(result.summary.length).toBeLessThan(500);
// 2. Expected Key Information Extraction
fixture.mustContainEntities.forEach((entity: string) => {
expect(result.summary.toLowerCase()).toContain(entity.toLowerCase());
});
// 3. Negative Constraint Check
fixture.forbiddenPhrases.forEach((phrase: string) => {
expect(result.summary.toLowerCase()).not.toContain(phrase.toLowerCase());
});
}, 15000); // 15s timeout for API call
});
});
3. Ragas Metrics for RAG Pipelines
When evaluating RAG systems, relying purely on lexical string matching (like BLEU or ROUGE) is ineffective. Modern pipelines employ Ragas metrics:
- Faithfulness: Quantifies whether every statement in the generated answer can be directly inferred from the retrieved context chunks (catches hallucinations).
- Answer Relevance: Measures whether the response directly addresses the user prompt rather than deflecting.
- Context Precision: Evaluates whether the vector retriever ranked relevant chunks higher than irrelevant noise.
# run_ragas_eval.py
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision
from datasets import Dataset
def evaluate_rag_pipeline(eval_data: dict):
dataset = Dataset.from_dict(eval_data)
results = evaluate(
dataset=dataset,
metrics=[faithfulness, answer_relevancy, context_precision]
)
print("Faithfulness Score:", results["faithfulness"])
assert results["faithfulness"] > 0.90, "Regression: Faithfulness fell below 90% threshold!"
4. Key Takeaways
- Never merge prompt changes without automated evals: Treat prompts with the same rigor as database migration scripts.
- Curate 100 Golden Fixtures: Store ground-truth test cases in JSON fixtures within version control.
- Fail Fast in CI: Run low-cost deterministic checks first, followed by tiered semantic evaluations on sample subsets.