现有AI监管依赖的测试基准无法反映真实表现,亟需新框架。
Beyond Benchmarks: On The False Promise of AI Regulation
- 不依赖测试基准,提出可直接落地的监管方案
- 实证指出基准成绩与实际风险无相关性
- 适合政策制定者和跨学科研究者参考
当前AI模型在安全测试基准上的表现,并不能反映其部署后的实际表现。这种模型内在的不可解释性,使基于基准成绩构建的监管体系失效,难以应对现实中的持续伤害。问题根源在于人工智能可解释性这一基础挑战,却常被监管者和决策者忽视。本文提出一种简单、现实且可立即应用的监管框架,不再依赖基准测试,并呼吁跨学科合作,探索解决这一关键问题的新路径。
原文摘要 · Abstract (English)
The performance of AI models on safety benchmarks does not indicate their real-world performance after deployment. This opaqueness of AI models impedes existing regulatory frameworks constituted on benchmark performance, leaving them incapable of mitigating ongoing real-world harm. The problem stems from a fundamental challenge in AI interpretability, which seems to be overlooked by regulators and decision makers. We propose a simple, realistic and readily usable regulatory framework which does not rely on benchmarks, and call for interdisciplinary collaboration to find new ways to address this crucial problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。