Reproducible AUC Evaluation Harness

testingChallenge

Prompt Content

Develop a standalone evaluation script or function that takes the trained model's weights and the test data, then computes the AUC score using standard libraries (e.g., scikit-learn). The harness should be robust, handle large datasets, and be capable of being executed reproducibly. Document any specific steps, configurations, or environment considerations needed for consistent and accurate AUC measurement.

Try this prompt

Open the workspace to execute this prompt with free credits, or use your own API keys for unlimited usage.

Usage Tips

Copy the prompt and paste it into your preferred AI tool (Claude, ChatGPT, Gemini)

Customize placeholder values with your specific requirements and context

For best results, provide clear examples and test different variations