要求大模型实验室公开小型可复现模型,兼顾安全与创新。
Position: Require Frontier AI Labs To Release Small "Analog" Models
- 要求大厂发布与主模型同源的小型开源模型。
- 小模型能有效验证安全性和可解释性,且效果可迁移至大模型。
- 降低监管成本,适合研究机构和公众参与模型安全研究。
当前前沿大模型的监管提案因安全与创新的权衡而难以推进。本文提出替代方案:强制大型AI实验室发布与其最大模型同源、经蒸馏训练的小型开源模型(即“模拟模型”)。这些模型作为公共代理,使更广泛的研究社区可在不接触核心模型的前提下,开展安全验证、可解释性研究和算法透明度探索。近期研究表明,基于小模型开发的安全与可解释性方法能有效泛化至前沿规模系统。该政策几乎不增加额外成本,可复用数据与基础设施资源,显著推动公共利益。我们希望此政策不仅被采纳,更能体现一个根本原则:对模型的深入理解可缓解安全与创新的矛盾,实现二者兼得。
原文摘要 · Abstract (English)
Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovation tradeoff. This paper argues for an alternative regulatory approach that ensures AI safety while actively promoting innovation: mandating that large AI laboratories release small, openly accessible analog models (scaled-down versions) trained similarly to and distilled from their largest proprietary models. Analog models serve as public proxies, allowing broad participation in safety verification, interpretability research, and algorithmic transparency without forcing labs to disclose their full-scale models. Recent research demonstrates that safety and interpretability methods developed using these smaller models generalize effectively to frontier-scale systems. By enabling the wider research community to directly investigate and innovate upon accessible analogs, our policy substantially reduces the regulatory burden and accelerates safety advancements. This mandate promises minimal additional costs, leveraging reusable resources like data and infrastructure, while significantly contributing to the public good. Our hope is not only that this policy be adopted, but that it illustrates a broader principle supporting fundamental research in machine learning: deeper understanding of models relaxes the safety-innovation tradeoff and lets us have more of both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。