要求顶会论文公开安全测试数据,否则不认可其结论。
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
- 提出三档披露框架,区分公开、受控和受限披露
- 指出当前模型安全评估平均透明度仅40分,普遍隐瞒训练测试重叠
- 适合关注AI治理与可复现性的研究者与政策制定者
前沿AI安全声明——即宣称高能力通用模型未达风险阈值、已充分缓解或适合发布——正日益影响模型部署、治理与公众信任。然而,用于评估这些声明的原始数据常被隐藏,导致证据倒置:最关键的AI安全主张往往最不可复现。本文主张NeurIPS应强制要求此类论文遵循可复现性标准,将非复现视为方法论失败而非透明度偏好。2026年国际AI安全报告指出,可靠的预发布安全测试已愈发困难,且模型能区分测试与部署环境;2025年基础模型透明度指数显示行业平均透明度仅为40/100,无主要开发者充分披露训练-测试重叠情况;同期测量理论研究发现,跨系统攻击成功率比较常基于低效度测量。本文提出三档披露框架,包括公开、受控与声明受限披露,并配套强制声明清单、范围说明及分阶段实施路径与渐进惩罚机制。该框架将保密与开放视为连续谱两端,受控审查(通过合格安全评审机构组成的联邦评议会)覆盖无法公开的声明,右扩展声明则连保密审查也无法进行。社区对最重要声明所适用的标准,不应低于对最不重要声明的标准。
原文摘要 · Abstract (English)
Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or suitable for release - increasingly shape model deployment, governance, and public trust. Yet the artefacts needed to evaluate them are routinely withheld, producing an evidential inversion: the most consequential claims in AI safety are often the least reproducible. This position paper argues that NeurIPS should require reproducibility standards for papers making such claims, treating non-reproducibility not as a transparency preference but as an evaluation-methodology failure. The 2026 International AI Safety Report [Bengio et al., 2026] concludes that reliable pre-deployment safety testing has become harder to conduct and that models now distinguish test from deployment contexts; the 2025 Foundation Model Transparency Index [Wan et al., 2025] reports a sector-average transparency score of 40/100 with no major developer adequately disclosing train-test overlap; contemporaneous measurement-theory work shows that attack-success-rate comparisons across systems are often founded on low-validity measurements [Chouldechova et al., 2025]. We propose a three-tier disclosure framework, distinguishing public, controlled, and claim-restricted disclosure, paired with a mandatory claim inventory, scope statements, and a phased implementation path with graduated sanctions. The framework treats secrecy and openness as endpoints of a spectrum, with controlled review (via a federated colloquium of qualified secure-review hosts) covering claims whose artefacts cannot be released publicly, and right-scaling claims whose artefacts cannot be reviewed even confidentially. The standard the community applies to its most consequential claims should be at least as high as the standard it applies to its least.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。