AI评估需公开假设,否则无法有效监管。
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
- 要求开发者明示并证明评估中的关键假设
- 若假设无法成立或评估发现重大风险,应暂停开发
- 提升AI研发透明度,助力高级AI治理
随着AI系统发展,AI评估已成为保障安全的重要监管支柱。我们主张,监管应要求开发者在论证安全性时,明确列出并合理说明评估中的核心假设。这些假设包括全面威胁建模、代理任务有效性以及充分的能力激发等,适用于现有模型评估与未来模型预测。目前许多假设尚缺乏充分依据。若监管以评估为基础,则当评估揭示不可接受的风险或假设无法合理证明时,应暂停开发。该方法旨在增强AI研发透明度,为高级AI的有效治理提供可行路径。
原文摘要 · Abstract (English)
As AI systems advance, AI evaluations are becoming an important pillar of regulations for ensuring safety. We argue that such regulation should require developers to explicitly identify and justify key underlying assumptions about evaluations as part of their case for safety. We identify core assumptions in AI evaluations (both for evaluating existing models and forecasting future models), such as comprehensive threat modeling, proxy task validity, and adequate capability elicitation. Many of these assumptions cannot currently be well justified. If regulation is to be based on evaluations, it should require that AI development be halted if evaluations demonstrate unacceptable danger or if these assumptions are inadequately justified. Our presented approach aims to enhance transparency in AI development, offering a practical path towards more effective governance of advanced AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。