🛢️ Data Engineering · Batch Pipelines
Test warehouse data like software (BigQuery)
Layered data quality tests that block bad data from reaching downstream consumers.
intermediate~30 minData EngineersAnalytics EngineersPlatform Data Teams
Steps
- 1Tier 0 volume/freshness: row counts within expected bands per partition
- 2Tier 1 structural: uniqueness of keys, nullability, referential integrity
- 3Tier 2 business invariants: revenue reconciles, statuses in enum sets
- 4Fail loudly at tier matching blast radius; warn below it
- 5Store test results historically to catch slow drift trends
- 6Review false-positive rate monthly; noisy tests get ignored
- 7Use partition pruning in test queries to control scan costs
- 8Assert on _TABLE_METADATA row counts cheaply
Common Pitfalls
- ▲Tests only in BI layer where nobody sees red
- ▲100% pass-rate suites proving nothing was tested
Commands
Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill data-quality-tests-warehouse-bigqueryInstall globally
$ npx skills add aniruddhaadak80/skills --skill data-quality-tests-warehouse-bigquery -gTags
#data-quality#testing#data-engineering#pipelines-batch