安全AI需处理不确定性与非阿基米德偏好。
Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities
- 引入不确定性推理与不完整偏好建模
- 支持非阿基米德效用以应对极端价值权衡
- 适合关注AI对齐与安全机制的研究者
如何确保AI系统与人类价值观对齐并保持安全?我们可通过AI协助和AI关机游戏框架来研究这一问题。AI协助问题涉及设计一个能帮助人类最大化其效用函数的智能体,但效用函数由人类知晓,AI助手必须学习它们。关机问题则要求设计出在按下关机按钮时能关闭、既不阻止也不诱使关机按钮被按下的智能体,并且在其他情况下仍能有效完成任务。本文表明,解决这些挑战需要能处理不确定性的智能体,并具备处理不完整及非阿基米德偏好(non-Archimedean preferences)的能力。
原文摘要 · Abstract (English)
How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down when a shutdown button is pressed; neither try to prevent nor cause the pressing of the shutdown button; and otherwise accomplish their task competently. In this paper, we show that addressing these challenges requires AI agents that can reason under uncertainty and handle both incomplete and non-Archimedean preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。