arXiv:2504.01849cs.AIcs.CY2025-04被引 52

提出技术方案防范AGI滥用与对齐失败带来的毁灭性风险

An Approach to Technical AGI Safety and Security

  • 识别四大风险,聚焦滥用与对齐问题,设计分级防御策略
  • 通过能力识别、访问控制和监控降低恶意使用可能性
  • 结合可解释性与不确定性估计,构建模型与系统双重安全保障

通用人工智能(AGI)虽具变革性潜力,但也带来足以严重危害人类的潜在风险。本文识别出四类风险:滥用、对齐失败、失误与结构性风险,重点探讨技术层面应对滥用与对齐问题的方案。针对滥用,提出主动识别危险能力,并实施严格安全措施、访问限制、监控及模型防护。针对对齐失败,构建双层防御:一是模型层面的增强监督与鲁棒训练;二是系统层面的监控与访问控制,即便模型失准也能减轻损害。可解释性、不确定性估计与更安全的设计模式可提升这些措施效果。最后简要说明如何整合各要素形成AGI系统的安全论证。

原文摘要 · Abstract (English)

Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough to significantly harm humanity. We identify four areas of risk: misuse, misalignment, mistakes, and structural risks. Of these, we focus on technical approaches to misuse and misalignment. For misuse, our strategy aims to prevent threat actors from accessing dangerous capabilities, by proactively identifying dangerous capabilities, and implementing robust security, access restrictions, monitoring, and model safety mitigations. To address misalignment, we outline two lines of defense. First, model-level mitigations such as amplified oversight and robust training can help to build an aligned model. Second, system-level security measures such as monitoring and access control can mitigate harm even if the model is misaligned. Techniques from interpretability, uncertainty estimation, and safer design patterns can enhance the effectiveness of these mitigations. Finally, we briefly outline how these ingredients could be combined to produce safety cases for AGI systems.

AGI安全对齐问题风险防范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。