arXiv:2503.17388cs.CYcs.AI2025-03被引 4

要求大模型公司公开安全评估前后数据,以便有效监管。

AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations

  • 要求企业同时报告模型部署前后的安全评估结果。
  • 当前评估多只做一种,易误导政策判断。
  • 适合关注AI监管与安全透明度的政策制定者。

人工智能系统的快速发展引发了对前沿AI系统潜在危害的广泛关注,亟需负责任的评估与监督。本文主张,前沿AI公司应报告模型部署前和部署后的安全评估结果,以支持科学决策。仅依赖任一阶段的评估会形成错误的安全图景。我们分析了领先实验室的安全披露发现三大问题:(1)极少同时评估部署前后版本;(2)评估方法缺乏统一标准;(3)报告结果普遍模糊不清。为此,我们建议强制向批准的政府机构披露前后评估结果,采用标准化评估方法,并设定公共安全报告的最低透明度要求,确保政策制定者能精准制定安全措施、评估部署风险并有效审查企业声明。

原文摘要 · Abstract (English)

The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this position paper, we argue that frontier AI companies should report both pre- and post-mitigation safety evaluations to enable informed policy decisions. Evaluating models at both stages provides policymakers with essential evidence to regulate deployment, access, and safety standards. We show that relying on either in isolation can create a misleading picture of model safety. Our analysis of AI safety disclosures from leading frontier labs identifies three critical gaps: (1) companies rarely evaluate both pre- and post-mitigation versions, (2) evaluation methods lack standardization, and (3) reported results are often too vague to inform policy. To address these issues, we recommend mandatory disclosure of pre- and post-mitigation capabilities to approved government bodies, standardized evaluation methods, and minimum transparency requirements for public safety reporting. These ensure that policymakers and regulators can craft targeted safety measures, assess deployment risks, and scrutinize companies' safety claims effectively.

AI安全监管透明度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。