⚡AgentSkills
🤖 AI Engineering · Inference & MLOps

Budget LLM latency end to end

Allocate milliseconds across retrieval, prompting, generation, and streaming so p95 meets product targets.

intermediate~30 minAI EngineersML EngineersLLM App Developers

Steps

  1. 1Trace one real request through every hop and record percentile timings
  2. 2Fix a p95 target from UX research, not infrastructure convenience
  3. 3Stream tokens to first-paint; perceived latency beats raw latency
  4. 4Cap retrieved context by latency cost, not just token count
  5. 5Cache stable prefixes: system prompts, few-shots, static context
  6. 6Alert when any hop exceeds 40% of total budget for an hour

Common Pitfalls

  • ▲Optimizing model choice while network hops dominate
  • ▲Non-streaming endpoints behind slow proxies

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill inference-mlops-latency-budgeting
Install globally
$ npx skills add aniruddhaadak80/skills --skill inference-mlops-latency-budgeting -g

Tags

#performance#latency#streaming#ai-engineering#inference-mlops

Related skills

Route by task complexity, cache aggressively, and enforce budgets so unit economics hold as usage grows.

🤖 AI Engineering·~30m

Treat all retrieved content as untrusted input: isolate instructions from data, gate actions, and fuzz continuously.

🤖 AI Engineering·~40m