Back to evaluations
Public evaluation

ma_codebase_audit_task

Evaluates accuracy of codebase scanning and licensing risk identification across fintech target code bases.

Evaluation type
task based
Challenge
Build an M&A Technical Due Diligence Agent with OpenAI Agents SDK
Difficulty
Intermediate
Rigor
Unspecified

Evaluation overview

How the linked challenge is judged: tasks, benchmarks, and criteria count.

Tasks
1
Benchmarks
0
Criteria
0

Task templates

Inputs and expected outputs.

Task 1

ma_codebase_audit_task

Analyze package specification file and return high-risk license flags and total technical debt score

Input format

JSON containing requirements_txt or package_json content

Output format

JSON with high_risk_licenses, cve_count, overall_risk_score, summary