Multi-Modal Edge AI for Defense with AutoGen and OpenAI o3/GPT-5
Challenge involves developing a sophisticated edge-to-cloud multi-modal agent system for real-time situational awareness. Participants will build a prototype where a simulated 'Edge Perception Agent' on a device processes visual and audio data, performing instant reasoning. This agent then securely communicates critical insights and requests for deeper analysis to a 'Cloud Command Agent' capable of advanced strategic reasoning and decision-making. The system must leverage AutoGen for orchestrating the heterogeneous agents and MCP for secure, efficient, and tool-integrated communication between the edge and cloud components. Focus on hybrid instant/deep reasoning, where quick local processing handles immediate threats, and a powerful cloud-based LLM is engaged for complex problem-solving based on aggregated or highly specific data.
What you are building
The core problem, expected build, and operating context for this challenge.
Challenge involves developing a sophisticated edge-to-cloud multi-modal agent system for real-time situational awareness. Participants will build a prototype where a simulated 'Edge Perception Agent' on a device processes visual and audio data, performing instant reasoning. This agent then securely communicates critical insights and requests for deeper analysis to a 'Cloud Command Agent' capable of advanced strategic reasoning and decision-making. The system must leverage AutoGen for orchestrating the heterogeneous agents and MCP for secure, efficient, and tool-integrated communication between the edge and cloud components. Focus on hybrid instant/deep reasoning, where quick local processing handles immediate threats, and a powerful cloud-based LLM is engaged for complex problem-solving based on aggregated or highly specific data.
Shared data for this challenge
Review public datasets and any private uploads tied to your build.
What you should walk away with
Master AutoGen for orchestrating heterogeneous multi-agent systems, including defining agent roles, capabilities, and inter-agent communication protocols.
Implement multi-modal data ingestion and preliminary processing at the simulated edge using a lightweight model like OpenAI o3 (conceptual for 2025's efficient edge models).
Design and build secure, MCP-enabled tool integration for agents to interact with simulated sensors, databases, and external APIs across the edge and cloud.
Deploy GPT-5 as the core reasoning engine for the 'Cloud Command Agent,' enabling complex strategic analysis and decision-making based on aggregated edge data.
Develop a hybrid instant/deep reasoning pattern, where instant decisions are made at the edge (e.g., basic object detection), and deeper context-aware analysis is offloaded to the cloud.
Build A2A protocol agents within AutoGen, ensuring robust communication for task handoffs and shared understanding between edge and cloud components.
[ok] Wrote CHALLENGE.md
[ok] Wrote .versalist.json
[ok] Wrote eval/examples.json
Requires VERSALIST_API_KEY. Works with any MCP-aware editor.
DocsAI Research & Mentorship
Participation status
You haven't started this challenge yet
Operating window
Key dates and the organization behind this challenge.
Find another challenge
Jump to a random challenge when you want a fresh benchmark or a different problem space.