⚡AgentSkills
🧮 ML Research Engineering · Evaluation Integrity

Choose benchmarks that mean something

Benchmark suites matched to target capabilities with contamination checks and honest scopes.

intermediate~30 minResearch EngineersML ScientistsPhD Researchers

Steps

  1. 1Define capability claims first; benchmarks follow claims, not reverse
  2. 2Check train/test contamination against your corpora before trusting numbers
  3. 3Prefer held-out private sets for anything heading to publication
  4. 4Report full suite including unfavorable tasks, not highlights
  5. 5State statistical significance; differences inside noise bands are ties
  6. 6Archive exact eval configs alongside results permanently

Common Pitfalls

  • ▲SOTA claimed on benchmarks the model effectively memorized
  • ▲Task suites drifting between model generations

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill evaluation-integrity-benchmark-selection-honesty
Install globally
$ npx skills add aniruddhaadak80/skills --skill evaluation-integrity-benchmark-selection-honesty -g

Tags

#benchmarks#evaluation#integrity#ml-research#evaluation-integrity

Related skills

Provenance, composition, collection process, and limitations recorded before modeling begins.

🧮 ML Research Engineering·~30m