AGI对齐可能带来误用风险,需权衡安全与滥用隐患。
Misalignment or misuse? The AGI alignment tradeoff
- 提出对齐与误用的双重风险权衡模型
- 多数现有对齐技术可能加剧人类滥用风险
- 强调治理与控制机制比技术对齐更关键
构建与人类目标对齐的系统被视为实现安全有益人工智能的核心路径。我们主张,未来通用智能(机器人)代理若未对齐将带来灾难性风险;但与此同时,对齐的AGI也极大增加人类恶意滥用的灾难性风险。尽管两者均严重且相互冲突,我们指出理论上存在不提升滥用风险的对齐方法。进一步实证分析表明,当前多数对齐技术及可预见改进,很可能加剧灾难性滥用风险。由于AI影响取决于社会背景,我们讨论了重要社会因素,并建议通过鲁棒性、AI控制方法以及尤其良好的治理来降低对齐AGI引发滥用灾难的风险。
原文摘要 · Abstract (English)
Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally intelligent (robotic) AI agents - poses catastrophic risks. At the same time, we support the view that aligned AGI creates a substantial risk of catastrophic misuse by humans. While both risks are severe and stand in tension with one another, we show that - in principle - there is room for alignment approaches which do not increase misuse risk. We then investigate how the tradeoff between misalignment and misuse looks empirically for different technical approaches to AI alignment. Here, we argue that many current alignment techniques and foreseeable improvements thereof plausibly increase risks of catastrophic misuse. Since the impacts of AI depend on the social context, we close by discussing important social factors and suggest that to reduce the risk of a misuse catastrophe due to aligned AGI, techniques such as robustness, AI control methods and especially good governance seem essential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。