Versalist helps developers test AI agents, inspect failures, and compare changes before release.
Start with one task
- Run the code-check example. Download all files and compare two versions with six fixed checks. No API key required.
- Connect your agent. Load a published challenge, record your command, and check its output.
- Understand your results. Find a failed check and decide what to change.
For supported hosted runs, start from a challenge. Read where your agent runs for runtime requirements.
Running a pilot for a team? Read team adoption for roles, effort, and what the pilot measures.
Evaluate agents
- Evaluation loop. Learn what each stage receives and produces.
- Challenges. Select, run, evaluate, or create a challenge environment.
- Sandboxes. Learn about managed execution, availability, account limits, cleanup, and usage.
- Tool catalog. Compare tools against task, infrastructure, and cost requirements.
- Skill bundles. Inspect and reuse versioned agent instructions.
Build with Versalist
- Coding agents. Configure OpenCode, Claude Code, Codex, Cursor, Pi, or Zed.
- CLI. Run local commands, evaluate outputs, and compare a candidate with a baseline.
- MCP tools. Eight tools on
versalist mcp. - vskill. Search, pull, and publish Skill Exchange bundles.
- Local model runs. Run a challenge with an open-weight model on your computer.
- API reference. Challenges, submissions, runs, environment versions, execution records, comparisons, and release decisions over HTTP.
- API keys. Create a key with only the required scopes.
- Integrations. Manage credentials for external model providers.
Manage Versalist
- Workspace. Find projects, prompts, tools, challenge records, and Vera tasks.
- Account. Manage profile data, security settings, billing, and account requests.
Resources
- Core terms. Check the meaning of Versalist and agent evaluation terms.
- Changelog. Read the current release notes.
- Frequently asked questions. Find answers and support routes.
- llms.txt. Load a plain-text documentation index into an agent or editor.