Each skill bundle packages a reusable agent behavior: a prompt, supporting files, and evaluation criteria. Browse the public catalog, review the full source, then install a private copy you can edit and experiment with.
109 published bundles ready to inspect and install
Build environments where agents write SQL, execute it, and get scored on result correctness
Formally specify UI tasks with clear start states, goal states, and evaluation criteria
Trade-offs between pixel-level interaction and DOM-level interaction for UI agents