Play 42
Computer Use Agent
Vision-based desktop and web automation replacing brittle RPA with screen understanding.
Vision-based desktop and web automation — AI agent controls applications via screenshots and mouse/keyboard, replacing brittle RPA with intelligent screen understanding for legacy systems and cross-app workflows. Runs securely in Azure Container Apps with action replay and rollback capabilities. Ideal for automating legacy enterprise software with no API surface.
Architecture Pattern
Vision-reasoning-action loop: screenshot capture, element detection, deterministic action execution
Azure Services
DevKit (.github Agentic OS)
- agent.md — root orchestrator with builder→reviewer→tuner handoffs
- 3 agents — Computer Use Builder (gpt-4o), Reviewer (gpt-4o-mini), Tuner (gpt-4o-mini)
- 3 skills — deploy (210 lines), evaluate (168 lines), tune (266 lines)
- 4 prompts — /deploy, /test, /review, /evaluate with agent routing
- .vscode/mcp.json — FrootAI MCP with OpenAI Vision + VM password inputs + envFile
TuneKit (AI Config)
- config/openai.json — gpt-4o vision, temp=0.1
- config/browser.json — resolution, timeouts, action limits
- config/guardrails.json — no credential entry, screenshot redaction
- evaluation/eval.py — Task completion >85%, Error rate <10%
Tuning Parameters
Machine evidence
FrootAI evidence lifecycle
This is an internal evidence maturity label, not third-party certification, accreditation, legal compliance, or a production guarantee. Missing or expired evidence demotes automatically; catalog claims cannot promote a play.
This play currently has design evidence only. A runnable scenario, endpoint evaluation, and build receipts are the next contiguous gates.
Repo Intelligence
v1A no-clone, revision-pinned map for agents and humans. Observed evidence is separated from inferred flow so the output stays useful without pretending to be a full call graph.