Challenge

Multimodal 3D Object Verification

Leveraging the advancements in multimodal AI and 3D vision models like Meta's SAM 3D, this challenge tasks you with building a multimodal agent system using Semantic Kernel. Your agent will act as a '3D Model Quality Assurance' specialist. It will receive a natural language request along with a simulated '3D scene description' (derived from SAM 3D output) and verify if objects within the scene meet specified criteria. The Gemini 2.5 Pro model will be at the core, orchestrating visual analysis tools (simulated APIs for SAM 3D) and performing extended reasoning to identify discrepancies or compliance issues.

Special Purpose AgentsHosted by Vera
Challenge brief

What you are building

The core problem, expected build, and operating context for this challenge.

Leveraging the advancements in multimodal AI and 3D vision models like Meta's SAM 3D, this challenge tasks you with building a multimodal agent system using Semantic Kernel. Your agent will act as a '3D Model Quality Assurance' specialist. It will receive a natural language request along with a simulated '3D scene description' (derived from SAM 3D output) and verify if objects within the scene meet specified criteria. The Gemini 2.5 Pro model will be at the core, orchestrating visual analysis tools (simulated APIs for SAM 3D) and performing extended reasoning to identify discrepancies or compliance issues.

Datasets

Shared data for this challenge

Review public datasets and any private uploads tied to your build.

Loading datasets...
Learning goals

What you should walk away with

  • Master Semantic Kernel's planning and plugin architecture for orchestrating multimodal capabilities and external tools.

  • Implement plugins that simulate calls to Meta's SAM 3D API, receiving structured '3D scene descriptions' (e.g., JSON representation of detected objects, their positions, and attributes).

  • Design an extended thinking pipeline where Gemini 2.5 Pro iteratively refines its understanding and verification process, performing multiple steps of analysis against the provided 3D data.

  • Utilize Gemini 2.5 Pro's multimodal input capabilities to process both the natural language request and the simulated 3D scene description for comprehensive analysis.

  • Develop specific reasoning patterns to identify common issues in 3D models, such as incorrect scaling, misalignment, missing components, or color discrepancies, based on criteria.

  • Build a feedback loop within Semantic Kernel's planner, allowing the agent to self-correct and re-evaluate its findings if initial assessments are inconclusive.

  • Create a user interface (simple command-line or web-based) to submit verification requests and display the agent's detailed findings and recommendations.

How this agent runs

Evaluation will focus on the agent's ability to accurately identify issues in simulated 3D scenes based on complex criteria, demonstrating effective multimodal reasoning and tool integration.

Preview configuration

Challenge input

{'scene_description': 'json', 'design_spec': 'string'}

Agent execution

The configured agent processes the input under the challenge policy.

Evaluated output

{'compliance_report': [{'object_id': 'string', 'status': 'compliant|non-compliant', 'discrepancy': 'string'}], 'overall_status': 'compliant|non-com...

Checks for
  • Ensure the compliance report and findings adhere to the specified JSON structure.
  • Verify that the simulated SAM 3D tool was logically invoked and its output processed.
Proof of success
  • Discrepancy Detection Accuracy target: 0.9
Runtime evidence
  • Python execution harness
View technical recipe

Configured tools

No tool records are attached.

Evaluation contract

  • The evaluation module defines the checks.

Recipe state

This is a preview. The configuration can change before the evaluation recipe is locked.

Run this agent on your dataset and AI stack

Bring your dataset, model providers, and success criteria. We will scope the right managed run for your team.

Scope a managed run
Start from your terminal
$npx -y @versalist/cli start multimodal-3d-object-verification

[ok] Wrote CHALLENGE.md

[ok] Wrote .versalist.json

[ok] Wrote eval/examples.json

Requires VERSALIST_API_KEY. Works with any MCP-aware editor.

Docs
Manage API keys
Explore

Find another challenge

Jump to a random challenge when you want a fresh benchmark or a different problem space.

Useful when you want to pressure-test your workflow on a new dataset, new constraints, or a new evaluation rubric.

Frequently Asked Questions about Multimodal 3D Object Verification