# Versalist documentation > Versalist helps developers test AI agents, inspect failures, and compare changes before release. A challenge contains a repeatable task, test cases, and scoring rules. Authorized trace capture records call metadata. It does not record payload text. ## Quick start for agents Requires Node.js 18.17 or later. list works without a key. start, submit, local challenge runs, MCP, and vskill require VERSALIST_API_KEY. 1. npx -y @versalist/cli list --search "" - Find a challenge slug. No API key. 2. Sign up at https://versalist.com/sign-in?mode=signup. An admin reviews every new account before it can be used. Once approved, create a key at https://versalist.com/profile/api-keys with read:challenges and submit:solutions, then export VERSALIST_API_KEY=vk_live_... 3. npm install -g @versalist/cli, then versalist start - Writes CHALLENGE.md, .versalist.json, and eval/examples.json when the challenge publishes public cases; deletes an existing eval/examples.json when it publishes none. Adds .versalist/ to .gitignore. 4. versalist run --command "" --label baseline - Writes run.json (command, exit status, duration, Git revision), stdout.log, and stderr.log to .versalist/runs//. Exits non-zero when the agent command fails or times out. 5. versalist evaluate --run latest --command "" - Score is 100 when the verifier exits 0 and 0 otherwise, unless you pass --score. Exits non-zero when the verifier fails. 6. versalist compare --baseline --candidate [--min-delta ] - Both runs must be evaluated with the same verifier command. decision is improved, unchanged, below_threshold, or regressed. Exits 1 unless the gate passes; usage errors also exit 1, so CI should read gate_passed from --json output. 7. Add --json to list, run, evaluate, or compare for machine-readable output. To use the tools inside an MCP host instead, run npx -y @versalist/cli mcp. Host configurations: https://versalist.com/docs/coding-agents ## Terminology User-facing term -> API term: run -> episode; execution record -> rollout; evaluation result or score -> reward; saved evaluation runs -> corpus; test cases and expected results -> gold items; proposed skill change -> amendment; review and release -> governance; saved failure -> finding. ## Documentation - [Documentation index](https://versalist.com/docs): Every guide grouped by task - [Quickstart](https://versalist.com/docs/getting-started): Complete local JavaScript code-check example with six fixed cases, no model or API key - [Run your agent](https://versalist.com/docs/run-your-agent): Connect an agent command to a published challenge, record local evaluations, and compare changes - [Trace capture](https://versalist.com/docs/trace-capture): Bounded call metadata, emitter support, configuration, and evidence limits - [Trust](https://versalist.com/trust): Verified product behavior and pending operational evidence - [Team adoption](https://versalist.com/docs/team-adoption): Pilot roles, set-up effort, workflow fit, and what a pilot measures - [Approval brief](https://versalist.com/enterprise/approval-brief): Browser-generated internal approval brief with labeled assumptions and attached review evidence - [Understand your results](https://versalist.com/docs/results): Scores, failed checks, execution errors, timeouts, traces, retention, access - [Test a change before you release it](https://versalist.com/docs/skills): Save a failure, propose a change, compare baseline and candidate, record the decision - [Where your agent runs](https://versalist.com/docs/where-your-agent-runs): Hosted sandbox, your own machine, and local CLI records compared - [Sandboxes](https://versalist.com/docs/sandboxes): Managed sandbox execution, availability, account settings, cleanup, usage, and billing - [Coding agents](https://versalist.com/docs/coding-agents): Select a coding agent and configure its CLI and MCP connection - [OpenCode setup](https://versalist.com/docs/coding-agents/opencode): Configure OpenCode with Versalist - [Claude Code setup](https://versalist.com/docs/coding-agents/claude-code): Configure Claude Code with Versalist - [Codex setup](https://versalist.com/docs/coding-agents/codex): Configure Codex with Versalist - [Cursor setup](https://versalist.com/docs/coding-agents/cursor): Configure Cursor with Versalist - [Pi setup](https://versalist.com/docs/coding-agents/pi): Configure Pi with Versalist - [Zed setup](https://versalist.com/docs/coding-agents/zed): Configure Zed with Versalist - [Evaluation loop](https://versalist.com/docs/agent-training-stack): Define a task, run your agent, inspect the results, and compare a change - [Challenges](https://versalist.com/docs/challenges): Select, run, evaluate, and create challenge environments - [Tool catalog](https://versalist.com/docs/ai-tools): Find tools and compare them against task requirements - [CLI reference](https://versalist.com/docs/cli): Load challenge files and record local evidence - [MCP tools](https://versalist.com/docs/mcp): Tools exposed by versalist mcp - [vskill](https://versalist.com/docs/vskill): Search, pull, and publish Skill Exchange bundles - [Local model runs](https://versalist.com/docs/local-model-runs): Run challenges with an open-weight model through Ollama - [API reference](https://versalist.com/docs/api): Challenges, submissions, runs, environment versions, execution records, comparisons, and release decisions over HTTP - [API keys](https://versalist.com/docs/api-keys): Create, scope, store, and revoke Versalist API keys - [Integrations](https://versalist.com/docs/integrations): Store and manage external provider credentials - [Workspace](https://versalist.com/docs/workspace): Find projects, prompts, tools, challenges, and Vera tasks - [Account](https://versalist.com/docs/account): Manage profile, security, billing, and account requests - [Core terms](https://versalist.com/docs/glossary): User-facing terms mapped to API names, plus definitions - [Changelog](https://versalist.com/docs/changelog): Recent product changes and release sources - [FAQ](https://versalist.com/docs/faq): Answers and support routes ## Public product routes - [Challenges](https://versalist.com/challenges): Challenge catalog - [Tools](https://versalist.com/ai-tools): Tool catalog - [Skill bundles](https://versalist.com/skill-bundles): Skill bundle catalog - [Prompt library](https://versalist.com/prompt-library): Public prompt library - [Evaluations](https://versalist.com/evaluations): Evaluation catalog - [Research](https://versalist.com/research): Research - [Guides](https://versalist.com/guides): Guides - [Changelog](https://versalist.com/changelog): Product release history ## Authenticated routes - /settings - Settings index (account and organization modules) - /profile - Profile data - /profile/settings - Preferences (default project visibility) - /profile/notifications - Email notification preferences - /profile/security - Sign-in providers and sessions - /profile/billing - Plan, invoice, credit, and customer portal actions - /profile/api-keys - Developer API: Versalist platform API keys and connected apps - /profile/integrations - Inference providers: external provider credentials and default provider - /company/settings - Organization general settings (admins) - /company/members - Organization members, invites, and roles (admins) - /workspace - Workspace index - /workspace/tools - Workspace tools - /workspace/stack - Stack definition - /workspace/vera - Vera Workbench - /episodes - Run history - /runs/{id} - Run detail: score by rubric dimension, task-by-task results, execution trace - /governance - Governance and corpus: qualified runs per week, execution queue, release-readiness gates (active company required) - /my-challenges - User challenge records - /my-projects - User projects - /my-prompts - User prompts - /progress/dashboard - Progress records - /progress/certificates - Certificate records ## Evaluation model 1. Define a challenge with inputs, constraints, outputs, and acceptance criteria. 2. Run your agent and record its evaluation result. 3. Apply deterministic checks, rubric criteria, baselines, or reviewer decisions. 4. Record the result, failure mode, and supporting evidence. 5. Change a skill only when the evidence supports the change. When trace capture is enabled, an Episode can include best-effort metadata for agent turns, agent model calls, judge calls, and (for sandbox runs) a sandbox_action event with status, exit code, duration, and byte counts. Traces do not include hidden model reasoning, external tool calls, or payload text. Eligible corpus records can retain separate payloads with consent. Public-case episode outputs and evaluator text are stored separately from this policy. Failure categories distinguish a wrong answer (rubric_shortfall, assertion_failed) from a run that could not execute (execution_error, runtime_error, timeout, provider_error, policy_denied, invalid_output). A skill change updates agent instructions. This workflow does not update model weights. ## Review and release workflow (governance) 1. Save a failed trace event as a finding with its original input (private test case). 2. Propose a skill change as an amendment (base version, proposed content, rationale). 3. Evaluate baseline (current version) and candidate (proposed version) on the same cases, environment version, model snapshot, and seed. 4. Register the comparison before the deferred sweeps execute; request a decision: promote | reject | manual_review (policy evaluation-decision-policy-v1: >= 5 cases, >= 3 points improvement, dimension regression <= 2, worst case regression <= 5, cost +20%, latency +25%). 5. Record a signed release: promote or rollback. Availability: registered comparisons, decisions, and releases are API-only. They use the active personal or company workspace and require trace-capture authorization, and an approved dated model snapshot. Hosted sweep execution returns 503 SECURITY_REVIEW_REQUIRED until the deployment passes external security review; release signing returns 503 SIGNING_UNAVAILABLE until configured. ## Where your agent runs - Platform execution: Versalist calls the model and evaluates the output. When available, sandbox challenges execute task programs in a temporary Python environment managed by Versalist. Programs use the standard library with no network access. Each task attempt uses a fresh sandbox. Versalist manages cleanup, which can continue after the run ends. - Your own machine: versalist challenge run with Ollama, or the run protocol at /api/v1/runs. Public cases only; provenance_level self_reported. - Local CLI records: versalist run / evaluate / compare write to .versalist/ and never upload. ## Sandboxes Sandbox execution is not generally available. Check the challenge runtime status before starting a run. Unsupported runtimes prevent the run from starting. When available, Sandbox execution in Settings controls new sandbox episodes and maximum task duration. Settings changes do not alter episodes that have already started. Cancellation stops further tasks. Work that has already started can take time to stop. Sandbox usage separates execution outcomes from cleanup status. An unavailable measurement does not mean zero usage. Where paid runs are available, review the displayed price and spending controls before starting a run. Use account billing information to check applicable charges. Custom containers, browser automation, persistent workspaces, and connections to your own sandbox service are unavailable. Read /docs/sandboxes for the public guide. ## API authentication Send a Versalist API key in the x-api-key header. Available user-key scopes: - read:challenges - submit:solutions - read:submissions - read:skills - write:skills - read:runs (run protocol reads; submit:solutions also works) - execute:runs (run protocol writes; submit:solutions also works) - read:governance (findings, amendments, comparisons, metrics) - write:governance (save findings, propose changes, register comparisons, decisions, releases, cancel background runs) Manage keys at /profile/api-keys. The key creation page does not yet list the four run and governance scopes. Environment: - VERSALIST_API_KEY - required for start, submit, local challenge runs, MCP, and vskill - VERSALIST_BASE_URL - optional, defaults to https://versalist.com - Authenticated CLI HTTP timeout is 20 seconds ## Challenge API - GET /api/challenges/public - Anonymous catalog (no key). limit default 20, max 100 - GET /api/challenges - Authenticated list. limit default 20, max 50. Scope: read:challenges - GET /api/challenges/{id-or-slug} - Challenge detail - GET /api/challenges/{id-or-slug}/markdown - Challenge brief - GET /api/challenges/{id-or-slug}/gold-items - Public reference items - GET /api/challenges/{id-or-slug}/leaderboard - Leaderboard entries - GET /api/challenges/{id}/submissions - Public submissions; user_only=true needs read:submissions - POST /api/challenges/submissions - Create a submission. challenge_id is a UUID. Scope: submit:solutions - GET /api/user/submissions - Key owner's submissions. Scope: read:submissions Duplicate submit returns 409. Challenge routes do not rate-limit (no 429 on this surface). ## Runs API - POST /api/episodes - Starts a hosted run. Requires session authentication and a UUID Idempotency-Key. - Required fields: challenge_id and skill_bundle_id. Optional fields: skill_bundle_version_id, model_id, comparison_role. - A baseline does not execute the selected skill. The request body limit is 65536 bytes. - Retry with the same request body and key. The server returns the accepted Episode and uses its saved execution snapshot. - A queued claim without an Episode becomes stale after five minutes. Retry the same request and key to claim it again. - Errors: 400 VALIDATION_ERROR, 403 RUN_POLICY_BLOCKED, 404 NOT_FOUND, 409 SANDBOX_RUNTIME_UNAVAILABLE, 409 AUTORESEARCH_ACTIVE, 409 IDEMPOTENCY_CONFLICT, 413 PAYLOAD_TOO_LARGE, 422 TOO_MANY_TEST_CASES, 429 EVALUATION_BUDGET_EXCEEDED, 500 INTERNAL_SERVER_ERROR, 503 EPISODE_REPLAY_LOOKUP_FAILED, 503 EPISODE_REPLAY_SNAPSHOT_UNAVAILABLE, 503 EPISODE_START_CLEANUP_FAILED, 503 EPISODE_START_RECOVERY_FAILED - GET /api/episodes - List runs (challenge_id, skill_bundle_id, limit 1-100, offset 0-1000000) - GET /api/episodes/{id} - Status (pending | running | completed | failed | cancelled), score_percentage, dimension_scores, steps, error_message (owner only) - POST /api/episodes/{id}/cancel - Cancel; returns { status: cancelled } or "Episode is not running" - GET /api/challenges/{id}/run-options - Runnable models and sandbox runtime status - GET|POST /api/v1/runs, GET /api/v1/runs/{id}, POST /api/v1/runs/{id}/cancel - Run protocol for your own runner (alias of /api/v1/local-runs). Scopes: read:runs / execute:runs or submit:solutions. Returns 404 when the protocol is disabled. ## Environment versions API - GET /api/v1/environments (read:challenges), POST (session) - Definitions - GET /api/v1/environments/{id} - Definition and versions - POST /api/v1/environments/{id}/versions (session) - Publish an immutable version from a blueprint (sandbox_type none | docker | browser, timeout_seconds 1-300) - POST /api/v1/environments/{id}/bindings (session) - Bind a version to a challenge - POST /api/v1/environments/runs (submit:solutions) - Enqueue a background run; 202 { job_id }; 503 SECURITY_REVIEW_REQUIRED while hosted execution is disabled; 429 RUN_LIMIT_EXCEEDED - GET /api/v1/environments/runs/{id} (read:challenges) - Background run status - POST /api/v1/environments/runs/{id}/cancel (write:governance) ## Execution records API (corpus) - GET /api/v1/corpus/rollouts, GET /api/v1/corpus/rollouts/{id} (read:challenges) - Records with rewards, payload_refs, episode_trace_summary - POST /api/v1/corpus/rollouts/{id}/replay (submit:solutions) - mode replay_verifiers | replay_full - POST /api/v1/corpus/sweeps (submit:solutions), GET /api/v1/corpus/sweeps/{id} - Up to 20 targets x 10 dated model snapshots, max 100 jobs - GET /api/v1/corpus/failures (read:challenges) - failure-v1 categories between from and to - GET|PUT /api/v1/corpus/payload-policy (session) - mode metadata_only | redacted | encrypted_raw, retention_days 1-30 (default metadata_only, 7) ## Governance API - GET|POST /api/v1/governance/findings - Save a failed trace event as a private test case - GET|POST /api/v1/governance/amendments - Proposed changes; GET /api/v1/governance/amendments/{id}/export for signed lineage - GET|POST /api/v1/governance/comparisons; POST /api/v1/governance/comparisons/{id}/decision - POST /api/v1/governance/releases - { decision_id, previous_release_id, action: promote | rollback } -> signed package - GET /api/v1/governance/metrics - Scopes: read:governance for GET, write:governance for POST. Session auth also accepted. ## MCP tools (@versalist/cli mcp) - list_challenges - requires read:challenges (unlike CLI list, this always uses a key) - get_challenge - get_challenge_markdown - get_evaluation_breakdown - get_gold_examples - get_leaderboard - submit_solution - requires submit:solutions - get_my_submissions - requires read:submissions ## Skill Exchange (@versalist/vskill) - CLI: search, pull, push, suggest, status, mcp - MCP: search_skills, get_skill, report_outcome - Scopes: read:skills, write:skills - HTTP prefix: /api/skills/registry ## Credential boundary Versalist API keys authenticate requests to Versalist. Provider credentials authorize supported calls through an external model provider. Manage provider credentials at /profile/integrations. Provider keys are available for OpenAI, Anthropic, Google Gemini, xAI, Mistral, Cohere, DeepSeek, AWS Bedrock, Azure OpenAI, Google Vertex AI, Cloudflare Workers AI, Groq, Cerebras, Together AI, Fireworks AI, Baseten, DeepInfra, Replicate (beta), OpenRouter, and Perplexity. Custom OpenAI-compatible endpoints (Ollama, vLLM, TGI, dedicated deployments) are available. Credentials are stored in Google Secret Manager, one secret per user and provider. General user-managed compute adapters are planned. ## Support - [Support answers](https://versalist.com/support/faq): Account and billing questions - [Feedback](https://versalist.com/feedback): Product defect and feedback form - [Request a demo](https://versalist.com/request-demo): Team evaluation and pilot scoping - support@versalist.com - Account, billing, or recovery requests - privacy@versalist.com - Privacy, data export, or deletion requests Source: https://versalist.com/docs