S
Claude CodeMeasure prompt and model changes with real metrics
LLM Eval Harness: Score Prompts Before You Ship
Setuproll editorial@setuproll90.0Overall score
A reproducible evaluation workflow that runs a test set against candidate models, grades answers with both code checks and an LLM judge, and tracks score deltas across versions. For anyone shipping an LLM feature who needs proof a prompt change actually helped instead of vibes.
90.0Score
5Components
Get this build
terminal
npx promptfoo@latest init && npx promptfoo evalWhat gets written
- CLAUDE.md
- .claude/agents/dataset-builder.md
- .claude/agents/judge-prompt-tuner.md
- .claude/agents/regression-reporter.md
- .mcp.json
Components
Model
- Claude Sonnet 5 (under test)
- Claude Opus 5 (judge)
- GPT-6 Sol
Stack
- promptfoo
- Inspect AI
- DuckDB
- pytest
MCP servers
- filesystem
- github
Subagents
- dataset-builder
- judge-prompt-tuner
- regression-reporter
How it works
- Define a golden test set with expected answers and rubrics
- Run every candidate model and prompt variant in one sweep
- Grade with exact-match plus an Opus 5 judge for open answers
- regression-reporter blocks the PR if win rate drops vs baseline
Summary
A reproducible evaluation workflow that runs a test set against candidate models, grades answers with both code checks and an LLM judge, and tracks score deltas across versions. For anyone shipping an LLM feature who needs proof a prompt change actually helped instead of vibes.
90.0 score
0 Reviews
Your rating
Sign in to post
Loading discussion...