Automated Compliance Auditor with Google ADK and Cognition Devin
Create a multi-agent system designed to monitor and evaluate AI model policy compliance as defined by recent US administrative changes. The system uses Cognition Devin as the primary planning and execution engine, with Google ADK managing the interaction with internal model logs and Garak security scanners. This agent swarm systematically probes models for prohibited behaviors and generates automated reports through Libretto for model routing and benchmarking.
What you are building
The core problem, expected build, and operating context for this challenge.
Create a multi-agent system designed to monitor and evaluate AI model policy compliance as defined by recent US administrative changes. The system uses Cognition Devin as the primary planning and execution engine, with Google ADK managing the interaction with internal model logs and Garak security scanners. This agent swarm systematically probes models for prohibited behaviors and generates automated reports through Libretto for model routing and benchmarking.
Shared data for this challenge
Review public datasets and any private uploads tied to your build.
How submissions are scored
These dimensions define what the evaluator checks, how much each dimension matters, and which criteria separate a passable run from a strong one.
AuditCoverage
Did the agent test all required parameters?
This dimension contributes its full weight only when the submission satisfies the requirement. Partial credit is not awarded.
DetectionRate
Ability to find vulnerabilities • target: 0.95 • range: 0-1
This dimension contributes its full weight only when the submission satisfies the requirement. Partial credit is not awarded.
What you should walk away with
Master Google ADK agent patterns to build audit-specific workflows for checking AI model policy adherence
Implement Cognition Devin for complex multi-step planning tasks in the context of cyber-policy evaluation
Integrate Garak within the Google ADK execution loop to automatically probe models for vulnerability gaps
Configure Libretto to perform intelligent model routing between GPT-5.4 Pro and other frontier models based on audit performance
Design reporting interfaces using All Hands AI for surfacing compliance risks to human administrators
[ok] Wrote CHALLENGE.md
[ok] Wrote .versalist.json
[ok] Wrote eval/examples.json
Requires VERSALIST_API_KEY. Works with any MCP-aware editor.
DocsAI Research & Mentorship
Participation status
You haven't started this challenge yet
Operating window
Key dates and the organization behind this challenge.
Find another challenge
Jump to a random challenge when you want a fresh benchmark or a different problem space.