Play 18
Prompt Management
Version control, A/B test, and rollback prompts across environments.
Manage prompts like code. Git-backed versioning, A/B testing with traffic splitting, automated quality scoring per version, rollback to any previous version. Cosmos DB stores prompt versions and experiment results. GitHub Actions CI/CD deploys prompt updates across dev/staging/prod.
Architecture Pattern
Prompt versioning, A/B testing, CI/CD, rollback, experimentation
Azure Services
DevKit (.github Agentic OS)
- agent.md — root orchestrator with builder→reviewer→tuner handoffs
- 3 agents — Prompt Mgmt Builder (gpt-4o), Reviewer (gpt-4o-mini), Tuner (gpt-4o-mini)
- 3 skills — deploy (111 lines), evaluate (101 lines), tune (114 lines)
- 4 prompts — /deploy, /test, /review, /evaluate with agent routing
- .vscode/mcp.json — FrootAI MCP with OpenAI + Cosmos DB inputs + envFile
TuneKit (AI Config)
- config/prompts.json — prompt versions, active versions
- config/ab-test.json — A/B weights, experiment tracking
- infra/ — Prompt Flow templates
Tuning Parameters
Machine evidence
FrootAI evidence lifecycle
This is an internal evidence maturity label, not third-party certification, accreditation, legal compliance, or a production guarantee. Missing or expired evidence demotes automatically; catalog claims cannot promote a play.
This play currently has design evidence only. A runnable scenario, endpoint evaluation, and build receipts are the next contiguous gates.
Repo Intelligence
v1A no-clone, revision-pinned map for agents and humans. Observed evidence is separated from inferred flow so the output stays useful without pretending to be a full call graph.