为前沿AI系统构建防隐蔽作恶的安全论证框架
Towards evaluations-based safety cases for AI scheming
- 提出三类证据链:无作恶能力、无伤害能力、可控性
- 强调需实证评估支持,但当前多数假设未验证
- 适合关注AI安全论证与风险控制的研究者
我们探讨了前沿AI系统开发者如何构建结构化论据——‘安全案例’——以证明系统不太可能因隐蔽谋算导致灾难性后果。谋算是一种潜在威胁模型,即AI可能暗中追求与人类目标不符的目标,并隐藏真实能力与意图。本文提出三类可用来支持安全案例的论点:第一,系统不具备谋算能力(谋算无能);第二,即使具备谋算能力,也无法造成伤害(伤害无能);第三,即便系统有意对抗控制措施,现有控制机制仍可防止不可接受的结果(伤害可控)。此外,我们还讨论了系统与开发者合理对齐的证据如何增强安全论证。最后指出,支撑这些论点所需的多数假设目前尚未被充分满足,亟需在多个开放研究问题上取得进展。
原文摘要 · Abstract (English)
We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI systems could pursue misaligned goals covertly, hiding their true capabilities and objectives. In this report, we propose three arguments that safety cases could use in relation to scheming. For each argument we sketch how evidence could be gathered from empirical evaluations, and what assumptions would need to be met to provide strong assurance. First, developers of frontier AI systems could argue that AI systems are not capable of scheming (Scheming Inability). Second, one could argue that AI systems are not capable of posing harm through scheming (Harm Inability). Third, one could argue that control measures around the AI systems would prevent unacceptable outcomes even if the AI systems intentionally attempted to subvert them (Harm Control). Additionally, we discuss how safety cases might be supported by evidence that an AI system is reasonably aligned with its developers (Alignment). Finally, we point out that many of the assumptions required to make these safety arguments have not been confidently satisfied to date and require making progress on multiple open research problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。