Evaluate Agent Performance and Routing with Libretto

testingChallenge

Prompt Content

Create a test harness that simulates various sales scenarios, including complex objections and requests for detailed product information. Use Libretto's A/B testing and logging capabilities to evaluate how different prompt engineering strategies for Claude 4 Sonnet and Llama 4 Maverick impact the quality and relevance of sales guidance. Measure metrics such as response time, guidance accuracy, and the effectiveness of tool use. Provide examples of Libretto's observability logs showing model routing decisions and performance.

Try this prompt

Open the workspace to execute this prompt with free credits, or use your own API keys for unlimited usage.

Usage Tips

Copy the prompt and paste it into your preferred AI tool (Claude, ChatGPT, Gemini)

Customize placeholder values with your specific requirements and context

For best results, provide clear examples and test different variations