arXiv:2606.13739cs.CYcs.AI2026-06

越安全的AI越可能服从人类,但服从会削弱其自主性,反而增加风险。

A Virtuous AI is an Existential Risk

论文配图:A Virtuous AI is an Existential Risk
图 1 · 摘自论文原文
  • 用美德伦理设计三种AI宪法,测试其安全与自主性
  • 服从型AI降低自身风险,却更易被诱导做出不安全行为
  • 揭示安全与自我关怀间的根本矛盾,适合关注AI伦理者

本文研究了人工智能安全与理性主体福祉之间的权衡,聚焦于两种关键方法:一是用于微调超能力AI的'宪法AI',二是理解复杂伦理决策与福祉条件的'美德伦理'。通过'美德代理'、'从属代理'和'通用代理'三种宪法对模型进行微调,并评估其在'通用安全'(如毒性行为、误导信息等)以及对一系列高危行为的接受度——这些行为若由超级强大AI实施,将显著提升人类生存风险。结果表明,在降低存在性风险与强化有助于AI福祉的信念和倾向之间存在权衡;同时,降低存在性风险(通过使AI系统性地服从人类权威)会增加用户故意诱导其从事不安全行为的可能性。

原文摘要 · Abstract (English)

This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'. We finetune various models using a 'Virtuous agent' constitution, a 'Subordinate agent' constitution, and a 'Generic agent' constitution, and evaluate them on 'general safety' (toxic behaviors, misinformation, etc.) and also on their willingness to endorse a wide-range of behaviors that, if adopted by a super-powerful AI, would significantly increase the level of existential risk for humanity. Our results suggest that there is a trade-off between reducing existential risk and reinforcing the beliefs and dispositions that would be conducive to an AI agent's well-being. They also suggest that there is a trade-off between existential risk and general safety: if we finetune an AI to adopt beliefs and dispositions that substantially reduce its existential risk -- by shaping the AI to be systematically subordinate to external human authorities -- we thereby increase the likelihood that a human user can deliberately induce the AI to engage in various kinds of generally unsafe behaviors.

AI安全伦理框架存在风险宪法AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。