现有评测基准严重偏离欧盟AI法案要求,无法评估关键系统风险。
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- 用大模型评分法系统分析19万+题目与法规能力分类的匹配度。
- 61.6%题量聚焦幻觉倾向,关键失控能力如自我复制零覆盖。
- 为监管合规评测工具研发提供首个量化依据,适合政策制定者参考。
通用人工智能(GPAI)模型的快速发展亟需稳健的评估框架,尤其在欧盟《人工智能法案》及其《行为准则》(CoP)出台背景下。当前的AI评估高度依赖传统基准,但这些工具未针对新法规关注的系统性风险设计。本研究解决“评测-监管”之间的核心差距,提出Bench-2-CoP框架,采用经验证的大语言模型作为评判者,系统分析194,955个来自主流基准的题目在欧盟AI法案能力与倾向分类体系中的覆盖率。结果揭示显著错位:评测体系将平均61.6%的监管相关问题集中在“幻觉倾向”,31.2%用于“性能不可靠性”,而关键功能性能力被严重忽视。尤其值得注意的是,涉及失控场景的核心能力——规避人类监督、自我复制及自主发展——在全部基准中均无覆盖。该研究首次提供了此差距的全面定量分析,表明现有公开基准单独不足以支持监管合规所需的风险评估证据,并为下一代评测工具开发提供关键洞见。
原文摘要 · Abstract (English)
The rapid advancement of General Purpose AI (GPAI) models necessitates robust evaluation frameworks, especially with emerging regulations like the EU AI Act and its associated Code of Practice (CoP). Current AI evaluation practices depend heavily on established benchmarks, but these tools were not designed to measure the systemic risks that are the focus of the new regulatory landscape. This research addresses the urgent need to quantify this "benchmark-regulation gap." We introduce Bench-2-CoP, a novel, systematic framework that uses validated LLM-as-judge analysis to map the coverage of 194,955 questions from widely-used benchmarks against the EU AI Act's taxonomy of model capabilities and propensities. Our findings reveal a profound misalignment: the evaluation ecosystem dedicates the vast majority of its focus to a narrow set of behavioral propensities. On average, benchmarks devote 61.6% of their regulatory-relevant questions to "Tendency to hallucinate" and 31.2% to "Lack of performance reliability", while critical functional capabilities are dangerously neglected. Crucially, capabilities central to loss-of-control scenarios, including evading human oversight, self-replication, and autonomous AI development, receive zero coverage in the entire benchmark corpus. This study provides the first comprehensive, quantitative analysis of this gap, demonstrating that current public benchmarks are insufficient, on their own, for providing the evidence of comprehensive risk assessment required for regulatory compliance and offering critical insights for the development of next-generation evaluation tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。