arXiv:2505.20203cs.AI2025-05被引 3

提出让智能体可关闭的新训练方法,避免其抗拒停机。

Shutdownable Agents through POST-Agency

  • 训练智能体只在相同长度轨迹间比较偏好
  • 理论证明能确保智能体不关心任务时长分布
  • 适合希望控制智能体、又需其高效执行的开发者

许多人担心未来的人工智能代理会抵抗关闭。本文提出一种名为 POST-Agents 的方案,即训练智能体仅在相同长度的轨迹之间进行偏好判断(POST)。随后证明,在其他条件下,POST 可推出 Neutrality+:智能体最大化期望效用,且忽略轨迹长度的概率分布。本文认为,Neutrality+ 能保证智能体可被关闭,同时仍保持有用性。

原文摘要 · Abstract (English)

Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I propose that we train agents to satisfy Preferences Only Between Same-Length Trajectories (POST). I then prove that POST - together with other conditions - implies Neutrality+: the agent maximizes expected utility, ignoring the probability distribution over trajectory-lengths. I argue that Neutrality+ keeps agents shutdownable and allows them to be useful.

智能体可控性偏好学习停机安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。