Versalist documentation

Test and improve your AI agents.

Versalist helps developers test AI agents, inspect failures, and compare changes before release.

Review one workflow with your team

Start with the team workflow and scoped pilot, then review execution and data flow, procurement status and team responsibilities. The published six-case example is local, self-reported sample code; a permissioned real agent comparison remains to be published.

Start with one task

  1. Run the code-check example. Download all files and compare two versions with six fixed checks. No API key required.
  2. Connect your agent. Load a published challenge, record your command, and check its output.
  3. Understand your results. Find a failed check and decide what to change.

New-account API keys require admin approval. This is separate from the no-key local example.

For supported hosted runs, start from a challenge. Read where your agent runs for runtime requirements.

Running a pilot for a team? Read team adoption for roles, effort, and what the pilot measures.

Use these docs from an AI agent

Point your coding agent at one of these entry points instead of scraping the pages.

  • /llms.txt: a plain-text index of every docs page, with a command quick start, API routes, scopes, and error codes.
  • npx -y @versalist/cli list --json: the public challenge catalog as JSON. No API key.
  • npx -y @versalist/cli mcp: an MCP server with eight challenge tools. It needs VERSALIST_API_KEY. Host setup is on coding agents.

Evaluate agents

  • Evaluation loop. Learn what each stage receives and produces.
  • Challenges. Select, run, evaluate, or create a challenge environment.
  • Where your agent runs. Compare the hosted sandbox, your own machine, and local CLI records.
  • Sandboxes. Learn about managed execution, availability, account limits, cleanup, and usage.
  • Trace capture. Check which call metadata is recorded and what it cannot prove.
  • Tool catalog. Compare tools against task, infrastructure, and cost requirements.
  • Compare skill changes. Save a failure, compare baseline and candidate, and record the decision.
  • Skill bundles. Inspect and reuse versioned agent instructions.

Build with Versalist

  • Coding agents. Configure OpenCode, Claude Code, Codex, Cursor, Pi, or Zed.
  • CLI. Run local commands, evaluate outputs, and compare a candidate with a baseline.
  • MCP tools. Eight tools on versalist mcp.
  • vskill. Search, pull, and publish Skill Exchange bundles.
  • Local model runs. Run a challenge with an open-weight model on your computer.
  • API reference. Challenges, submissions, runs, environment versions, execution records, comparisons, and release decisions over HTTP.
  • API keys. Create a key with only the required scopes.
  • Integrations. Manage credentials for external model providers.

Manage Versalist

  • Workspace. Find projects, prompts, tools, challenge records, and Vera tasks.
  • Account. Manage profile data, security settings, billing, and account requests.
  • Team adoption. Plan pilot roles, set-up effort, and what the pilot measures.

Resources

Was this page helpful?