Integrity Check and Reporting

testingChallenge

Prompt Content

Develop the final integrity check and reporting module. This module should analyze the results from your DSPy evaluation pipeline and identify patterns that suggest 'fudged' performance (e.g., disproportionate scores across benchmark subsets). Generate a comprehensive, auditable report detailing the findings, confidence levels, and the reasoning trace.

Try this prompt

Open the workspace to execute this prompt with free credits, or use your own API keys for unlimited usage.

Usage Tips

Copy the prompt and paste it into your preferred AI tool (Claude, ChatGPT, Gemini)

Customize placeholder values with your specific requirements and context

For best results, provide clear examples and test different variations