🤖 AI Engineering · RAG Pipelines
Evaluate RAG answer quality automatically
Score groundedness, relevance, and completeness with judge prompts plus deterministic checks wired into CI.
advanced~45 minAI EngineersML EngineersLLM App Developers
Steps
- 1Freeze 30-100 test questions with known-good source passages
- 2Run deterministic checks first: citation present, answer not empty, no refusal
- 3Use an LLM judge with rubric scores 1-5 for groundedness and relevance
- 4Swap judge models occasionally; if two judges disagree widely, fix the rubric
- 5Log every eval artifact: question, chunks used, answer, scores, cost, latency
- 6Block merges when groundedness average drops more than 0.3 below baseline
Common Pitfalls
- ▲Judging with the same model that generated the answer, inflating scores
- ▲Evals only run manually, drifting from production reality
Success Signals
- ✓Groundedness average 4.3+ of 5
- ✓Eval suite runtime under 10 minutes
Commands
Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill rag-pipelines-eval-rag-qualityInstall globally
$ npx skills add aniruddhaadak80/skills --skill rag-pipelines-eval-rag-quality -gTags
#evals#quality#ci#ai-engineering#rag-pipelines