Python sandbox execution is implemented. Hosted validation is pending. Availability depends on the service and the challenge requirements. If the required runtime is unavailable, the run cannot start. The challenge page explains the runtime status. Read the sandbox guide for account settings, limits, cleanup, and usage.
The three options
| Platform execution | Your own machine | Local CLI records | |
|---|---|---|---|
| What runs where | Versalist calls the model and evaluates the output. When available, sandbox challenges execute the generated Python program in a managed sandbox. | Your computer runs the model (through Ollama) or your agent. Versalist judges the returned answers and stores the run. | Your computer runs everything. The CLI writes a record of the command, its output, and your verifier result to .versalist/. |
| Start it with | Run on a challenge page, or POST /api/episodes | versalist challenge run … --model ollama:<tag>, or the run protocol at /api/v1/runs | versalist run, versalist evaluate, versalist compare |
| Prerequisites | A signed-in account. Sandbox challenges also require an available runtime and compatible account limits. | A Versalist API key, Ollama running with a model installed, and a challenge that allows local execution. | Node.js 18 or later and an API key for versalist start. |
| Cases evaluated | Public and private cases (evaluation_scope: full_suite) | Public cases only (evaluation_scope: public_suite) | Whatever your verifier checks. Public cases are in eval/examples.json. |
| Verification status | provenance_level: platform_verified. Versalist made the model call and can attest to the model used. | provenance_level: self_reported. The model identity is your claim; the run is labelled as local execution and stays off the trusted leaderboard. | None. The record is evidence for you and your reviewers, not a platform-verified result. |
| Where the result lives | Run history (/episodes) and the run page | Run history, marked as local execution | .versalist/runs/ and .versalist/comparisons/ in your repository |
Platform execution and the hosted sandbox
Model-only challenges evaluate model output directly. Some challenges require a program to process test inputs. For these challenges, Versalist evaluates the output from a sandbox execution.
Supported runtimes
The managed runtime supports restricted Python execution when available. Custom containers, browser automation, and connections to your own sandbox service are unavailable. A challenge must fit the supported runtime before a run can start.
What the Python sandbox allows
- Python standard library only. Package installation is unavailable.
- No network access or mounted files. Programs read task input and return output.
- A fresh sandbox for each task attempt. Files and running processes do not carry over between attempts.
- Time and output limits for every task.
How availability is shown
The challenge run card shows the required runtime and its availability. If the runtime is unavailable, the Run button is disabled and the card explains why. Account limits can also prevent a new sandbox episode from starting.
Cancellation
Cancel a running hosted episode from its run page. Cancellation stops further tasks. Work that has already started can take time to stop. Versalist manages sandbox cleanup.
Execution status and cleanup status are separate. Cleanup can continue after the run ends. See Cancellation and cleanup for more information.
Execution on your own machine
Use this to run an open-weight model on your own computer. Case responses leave your computer for platform evaluation. The CLI sends public cases to Ollama on your computer, returns the answers, and Versalist judges them. Setup and troubleshooting are on Local model runs.
- The platform records the run as local execution with a self-reported model.
- Only public cases are evaluated, so the score is not comparable to a full-suite hosted run.
- The challenge owner must allow local execution. The CLI cannot override that.
- The underlying run protocol at
/api/v1/runsalso acceptsopenai,anthropic,google-ai, andopenai-compatibleproviders for your own runner. Accepting a provider name does not mean Versalist hosts that provider.
Local CLI records
versalist run, evaluate, and compare never call Versalist. They write JSON records and logs under .versalist/ and exit non-zero when a check fails, which makes them usable as a CI gate. They are evidence for your own review; the platform does not verify them and submit does not upload them.
Fixed test environments
A hosted run with an enabled environment binding records an environment version: an immutable, digested description of the runtime, verifier, and evaluator the challenge was bound to when the run started. A shared environment version is one comparison requirement. Also match the cases, model, and generation settings.
Environment definitions and versions are managed through /api/v1/environments. Publishing a version and binding it to a challenge require a signed-in session. Executing a run against a published version through POST /api/v1/environments/runs is gated separately and returns 503 SECURITY_REVIEW_REQUIRED while hosted execution for that path is disabled.
Sandbox management
Versalist manages sandbox creation, execution, and cleanup. When available, account settings control new sandbox episodes and maximum task duration. Read Account settings for these controls.