Play 14
Cost-Optimized AI Gateway
APIM-based AI gateway with semantic caching, token budgets, and load balancing.
Route AI requests through APIM with semantic caching (Redis stores embeddings of recent queries — similar questions get cached responses). Token budgets per tenant prevent runaway costs. Multi-region load balancing with fallback chains ensures availability. Built-in analytics track cost per team.
Architecture Pattern
Semantic caching, token metering, load balancing, FinOps
Azure Services
DevKit (.github Agentic OS)
- agent.md — root orchestrator with builder→reviewer→tuner handoffs
- 3 agents — Gateway Builder (gpt-4o), Reviewer (gpt-4o-mini), Tuner (gpt-4o-mini)
- 3 skills — deploy (120 lines), evaluate (101 lines), tune (116 lines)
- 4 prompts — /deploy, /test, /review, /evaluate with agent routing
- .vscode/mcp.json — FrootAI MCP with APIM + subscription key inputs + envFile
TuneKit (AI Config)
- config/gateway.json — caching rules, token budgets, fallback chains
- config/routing.json — load balancing, model selection
- config/pricing.json — cost limits per tenant
Tuning Parameters
Machine evidence
Certified runtime lifecycle
Maturity is calculated from contiguous, content-bound evidence. Missing or expired evidence demotes automatically; catalog claims cannot promote a play.
This play currently has design evidence only. A runnable scenario, endpoint evaluation, and build receipts are the next contiguous gates.
Repo Intelligence
v1A no-clone, revision-pinned map for agents and humans. Observed evidence is separated from inferred flow so the output stays useful without pretending to be a full call graph.