Play 96
Real-Time Voice Agent v2
Next-gen bidirectional voice agent with sub-200ms latency and MCP tool integration.
Next-generation bidirectional WebSocket voice agent using Azure AI Voice Live SDK with MCP tool integration, voice activity detection, function calling during conversation, avatar rendering, and real-time transcription. Sub-200ms response latency for natural conversational AI.
Architecture Pattern
Voice agent loop: audio capture - VAD - STT - LLM reasoning - function calling - TTS - avatar rendering - transcription
Azure Services
DevKit (.github Agentic OS)
- agent.md — root orchestrator with builder→reviewer→tuner handoffs
- 3 agents — Voice V2 Builder (gpt-4o), Reviewer (gpt-4o-mini), Tuner (gpt-4o-mini)
- 3 skills — deploy (248 lines), evaluate (120 lines), tune (235 lines)
- 4 prompts — /deploy, /test, /review, /evaluate with agent routing
- .vscode/mcp.json — FrootAI MCP with OpenAI + Speech inputs + envFile
TuneKit (AI Config)
- config/openai.json - conversational prompts and function schemas
- config/voice.json - VAD mode, latency targets, avatar quality
- config/guardrails.json - response latency SLA, safety thresholds
- evaluation/eval.py - Latency <200ms P95, User satisfaction >4.2
Tuning Parameters
Machine evidence
FrootAI evidence lifecycle
This is an internal evidence maturity label, not third-party certification, accreditation, legal compliance, or a production guarantee. Missing or expired evidence demotes automatically; catalog claims cannot promote a play.
This play currently has design evidence only. A runnable scenario, endpoint evaluation, and build receipts are the next contiguous gates.
Repo Intelligence
v1A no-clone, revision-pinned map for agents and humans. Observed evidence is separated from inferred flow so the output stays useful without pretending to be a full call graph.