⚡AgentSkills
🤖 AI Engineering · RAG Pipelines

Evaluate RAG answer quality automatically

Score groundedness, relevance, and completeness with judge prompts plus deterministic checks wired into CI.

advanced~45 minAI EngineersML EngineersLLM App Developers

Steps

  1. 1Freeze 30-100 test questions with known-good source passages
  2. 2Run deterministic checks first: citation present, answer not empty, no refusal
  3. 3Use an LLM judge with rubric scores 1-5 for groundedness and relevance
  4. 4Swap judge models occasionally; if two judges disagree widely, fix the rubric
  5. 5Log every eval artifact: question, chunks used, answer, scores, cost, latency
  6. 6Block merges when groundedness average drops more than 0.3 below baseline

Common Pitfalls

  • ▲Judging with the same model that generated the answer, inflating scores
  • ▲Evals only run manually, drifting from production reality

Success Signals

  • ✓Groundedness average 4.3+ of 5
  • ✓Eval suite runtime under 10 minutes

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill rag-pipelines-eval-rag-quality
Install globally
$ npx skills add aniruddhaadak80/skills --skill rag-pipelines-eval-rag-quality -g

Tags

#evals#quality#ci#ai-engineering#rag-pipelines

Related skills

Choose chunk sizes, overlaps, and structure-aware splits so retrieved context actually helps the model answer.

🤖 AI Engineering·~25m

Combine BM25-style lexical search with dense embeddings and fuse results so both rare terms and paraphrases are found.

🤖 AI Engineering·~35m