🤖 AI Engineering · Inference & MLOps
Budget LLM latency end to end
Allocate milliseconds across retrieval, prompting, generation, and streaming so p95 meets product targets.
intermediate~30 minAI EngineersML EngineersLLM App Developers
Steps
- 1Trace one real request through every hop and record percentile timings
- 2Fix a p95 target from UX research, not infrastructure convenience
- 3Stream tokens to first-paint; perceived latency beats raw latency
- 4Cap retrieved context by latency cost, not just token count
- 5Cache stable prefixes: system prompts, few-shots, static context
- 6Alert when any hop exceeds 40% of total budget for an hour
Common Pitfalls
- ▲Optimizing model choice while network hops dominate
- ▲Non-streaming endpoints behind slow proxies
Commands
Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill inference-mlops-latency-budgetingInstall globally
$ npx skills add aniruddhaadak80/skills --skill inference-mlops-latency-budgeting -gTags
#performance#latency#streaming#ai-engineering#inference-mlops