ArgumentGeneration_ProTax
Evaluation will focus on the quality, factual accuracy, logical coherence, and persuasive strength of the arguments generated by the agent system for specified policy scenarios. Gentrace will be used to track and score these metrics.
Evaluation overview
How the linked challenge is judged: tasks, benchmarks, and criteria count.
Task templates
Inputs and expected outputs.
ArgumentGeneration_ProTax
Generates arguments in favor of an 'AI tax' on synthetic actors, given a set of policy documents.
{"stakeholder": "SAG-AFTRA", "topic": "AI Tax on Synthetic Actors", "documents": [{"title": "string", "content": "string"}]}
{"argument_text": "string", "key_points": ["string"], "potential_impacts": ["string"]}
ArgumentGeneration_AntiTax
Generates counter-arguments against an 'AI tax' on synthetic actors, given the same policy documents.
{"stakeholder": "Studio Executives", "topic": "AI Tax on Synthetic Actors", "documents": [{"title": "string", "content": "string"}]}
{"argument_text": "string", "key_points": ["string"], "potential_impacts": ["string"]}