⚡AgentSkills
🛢️ Data Engineering · Batch Pipelines

Make batch pipelines rerunnable by design

Idempotent stages, partitioned writes, and backfill strategies so failures heal instead of corrupt.

intermediate~35 minData EngineersAnalytics EngineersPlatform Data Teams

Steps

  1. 1Key every write by logical date/partition; overwrite partitions, never append blind
  2. 2Design stages to be safely re-executable from any point of failure
  3. 3Separate extraction, transformation, loading with checkpoints between
  4. 4Backfill = same code, different parameter — prove it in staging first
  5. 5Emit per-run metrics: rows in/out, duration, null-rates into a meta table
  6. 6Alert on freshness SLA breach before stakeholders notice stale dashboards

Common Pitfalls

  • ▲append-only loads duplicating on retry
  • ▲One giant DAG task failing everything for one bad record

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill pipelines-batch-pipeline-idempotency
Install globally
$ npx skills add aniruddhaadak80/skills --skill pipelines-batch-pipeline-idempotency -g

Tags

#etl#idempotency#orchestration#data-engineering#pipelines-batch

Related skills

Schema contracts, compatibility modes, and deprecation flows across pipeline boundaries.

🛢️ Data Engineering·~40m

Layered data quality tests that block bad data from reaching downstream consumers.

🛢️ Data Engineering·~30m

Layered data quality tests that block bad data from reaching downstream consumers.

🛢️ Data Engineering·~30m