ML Research Engineering
Rigorous ML experimentation: reproductions, ablations, tracking, and honest benchmarking.
Playbook: Ground scaling decisions in reality
Scaling-law claims converted into budget-honest training choices.
Experimentation Rigor
Staged reproduction from inference-first to full retrain with deviation journals.
One-factor-at-a-time with matched budgets separating real gains from tuning luck.
Every run reproducible from logged config + code version + data snapshot reference.
Evaluation Integrity
Benchmark suites matched to target capabilities with contamination checks and honest scopes.
Provenance, composition, collection process, and limitations recorded before modeling begins.
Inference Optimization
Drafter selection, acceptance-rate monitoring, and batch-interaction effects in production serving.
Compute budgets, data-wall caveats, and extrapolation ranges separating signal from slide-ware.
Journey Playbooks
Scaling-law claims converted into budget-honest training choices.