⚡AgentSkills
📊 Data Science & Analytics · Experimentation & A/B Testing

Design an A/B test you can trust

Power the test upfront, guard metrics, and commit to decision rules before peeking.

intermediate~35 minData ScientistsAnalystsResearch Scientists

Steps

  1. 1Define one primary metric with minimum detectable effect from business value
  2. 2Compute required sample size for 80% power at α=0.05 two-sided
  3. 3Randomize at user level with sticky assignment across sessions
  4. 4Add guardrail metrics: latency, error rate, unsubscribe, revenue per user
  5. 5Commit to runtime and decision rule BEFORE launch; no mid-flight goalposts
  6. 6Log the design doc; results include confidence intervals not just p-values

Common Pitfalls

  • ▲Peeking daily and stopping at first significance
  • ▲Underpowered tests 'proving' null effects

Success Signals

  • ✓100% of launched tests with pre-registered designs
  • ✓SRM check passing (sample ratio mismatch)

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill experimentation-ab-test-design
Install globally
$ npx skills add aniruddhaadak80/skills --skill experimentation-ab-test-design -g

Tags

#ab-testing#statistics#product#data-science#experimentation

Related skills

One canonical definition per metric, versioned and documented, killing dashboard wars.

📊 Data Science & Analytics·~30m