⚡AgentSkills
🤖 AI Engineering · Inference & MLOps

Control LLM spend without killing quality

Route by task complexity, cache aggressively, and enforce budgets so unit economics hold as usage grows.

intermediate~30 minAI EngineersML EngineersLLM App Developers

Steps

  1. 1Tag every request with feature and tenant for per-unit cost attribution
  2. 2Route simple tasks to small models using a trained classifier or heuristics
  3. 3Cache exact and semantic matches; measure hit-rate weekly
  4. 4Compress context: prune, summarize history, cap tool outputs
  5. 5Set per-key and per-tenant daily caps with graceful degradation paths
  6. 6Review the top ten most expensive prompts monthly and shrink them

Common Pitfalls

  • ▲Premium model answering 'is this email spam?'
  • ▲Semantic caches keyed too loosely, returning stale answers

Commands

Install with skills CLI
$ npx skills add aniruddhaadak80/skills --skill inference-mlops-llm-cost-controls
Install globally
$ npx skills add aniruddhaadak80/skills --skill inference-mlops-llm-cost-controls -g

Tags

#cost#routing#caching#ai-engineering#inference-mlops

Related skills

Allocate milliseconds across retrieval, prompting, generation, and streaming so p95 meets product targets.

🤖 AI Engineering·~30m

Treat all retrieved content as untrusted input: isolate instructions from data, gate actions, and fuzz continuously.

🤖 AI Engineering·~40m