Evaluation and Refinement

testingChallenge

Prompt Content

Run your complete agent system with the sample input for the 'Personalized Investment Query' task. Capture the voice response transcript and the JSON summary. Critically evaluate the agent's performance in terms of accuracy, personalization, and conversational flow. Suggest potential improvements, especially regarding how the agent handles ambiguous queries or multiple tracked entities. Provide the final runnable Python code.

Try this prompt

Open the workspace to execute this prompt with free credits, or use your own API keys for unlimited usage.

Usage Tips

Copy the prompt and paste it into your preferred AI tool (Claude, ChatGPT, Gemini)

Customize placeholder values with your specific requirements and context

For best results, provide clear examples and test different variations