⚡AgentSkills
🛢️ Data Engineering · Batch Pipelines

Test warehouse data like software (BigQuery)

Layered data quality tests that block bad data from reaching downstream consumers.

intermediate~30 minData EngineersAnalytics EngineersPlatform Data Teams

Steps

  1. 1Tier 0 volume/freshness: row counts within expected bands per partition
  2. 2Tier 1 structural: uniqueness of keys, nullability, referential integrity
  3. 3Tier 2 business invariants: revenue reconciles, statuses in enum sets
  4. 4Fail loudly at tier matching blast radius; warn below it
  5. 5Store test results historically to catch slow drift trends
  6. 6Review false-positive rate monthly; noisy tests get ignored
  7. 7Use partition pruning in test queries to control scan costs
  8. 8Assert on _TABLE_METADATA row counts cheaply

Common Pitfalls

  • ▲Tests only in BI layer where nobody sees red
  • ▲100% pass-rate suites proving nothing was tested

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill data-quality-tests-warehouse-bigquery
Install globally
$ npx skills add aniruddhaadak80/skills --skill data-quality-tests-warehouse-bigquery -g

Tags

#data-quality#testing#data-engineering#pipelines-batch

Related skills

Idempotent stages, partitioned writes, and backfill strategies so failures heal instead of corrupt.

🛢️ Data Engineering·~35m

Schema contracts, compatibility modes, and deprecation flows across pipeline boundaries.

🛢️ Data Engineering·~40m

Layered data quality tests that block bad data from reaching downstream consumers.

🛢️ Data Engineering·~30m