🛢️ Data Engineering · Lakehouse & Streaming
Choose and operate open table formats deliberately
Iceberg/Delta/Hudi tradeoffs, compaction strategy, and catalog governance preventing metadata debt.
advanced~40 minData EngineersAnalytics EngineersPlatform Data Teams
Steps
- 1Select by engine ecosystem fit and vendor-neutral catalog support first
- 2Define compaction/optimize cadence per table by write pattern
- 3Set snapshot retention policies balancing time-travel needs vs storage burn
- 4Govern the catalog: naming, ownership, staging-to-prod promotion flows
- 5Test schema evolution paths actually used (add column vs type widening)
- 6Monitor metadata file counts — small-file problems metastasize into query planning
Common Pitfalls
- ▲Every-write commits leaving thousands of tiny metadata files
- ▲Time-travel retention zero making recovery impossible
Commands
Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill lakehouse-streaming-open-table-format-selectionInstall globally
$ npx skills add aniruddhaadak80/skills --skill lakehouse-streaming-open-table-format-selection -gTags
#iceberg#delta-lake#lakehouse#data-engineering#lakehouse-streaming