Implement Evaluation with Continue.dev

testingChallenge

Prompt Content

Integrate Continue.dev to create an iterative testing loop for your proof assistant. Define test cases based on the provided `eval_data.json` (which contains new mathematical statements and their expected proofs/counter-examples). Use Continue.dev to automatically run tests against your LlamaIndex agent, compare generated proofs against ground truth, and report discrepancies. Focus on how Continue.dev helps in rapid iteration and debugging of agent behavior.

Try this prompt

Open the workspace to execute this prompt with free credits, or use your own API keys for unlimited usage.

Usage Tips

Copy the prompt and paste it into your preferred AI tool (Claude, ChatGPT, Gemini)

Customize placeholder values with your specific requirements and context

For best results, provide clear examples and test different variations