📊 Data Science & Analytics · Modeling Practice
Prevent data leakage in ML pipelines
Split-aware preprocessing, temporal boundaries, and feature provenance keeping offline gains real online.
advanced~35 minData ScientistsAnalystsResearch Scientists
Steps
- 1Split temporally when production predicts the future
- 2Fit scalers/imputers/encoders inside folds only, never full data
- 3Ban features unavailable at prediction timestamp (audit with data team)
- 4Group correlated rows (same user/device) into same split side
- 5Track feature computation code versions with model artifacts
- 6Validate: offline-vs-online correlation checked each release
Common Pitfalls
- ▲Random shuffles leaking user histories across train/test
- ▲Target encoding fit on full dataset
Commands
Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill ml-modeling-leakage-preventionInstall globally
$ npx skills add aniruddhaadak80/skills --skill ml-modeling-leakage-prevention -gTags
#ml#validation#leakage#data-science#ml-modeling