arXiv:2505.20342cs.AI2025-05中稿 · NeurIPS

用心理模型从行为推断人类价值观,解决有限样本下的价值泛化难题。

Machine Theory of Mind and the Structure of Human Values

  • 基于贝叶斯理论心智,从行为和其他价值观推断未知价值
  • 人类价值观具有生成性理性结构,可支持跨值推理
  • 适合研究可扩展机器心智与安全对齐的学者

价值学习是实现安全且合乎伦理人工智能的关键。当前主要方法是通过行为推断人类价值观。然而,人类关心的事物远超其行为所展现的内容。因此,人工智能必须从有限的行为样本中预测出剩余的复杂价值观。我将此称为价值泛化问题。本文主张,人类价值观具有生成性理性结构,这一特性使价值泛化成为可能。具体而言,我们可使用贝叶斯理论心智模型,不仅从行为,还能从其他价值观中推断人类价值。这一可能性长期被简单效用函数对人类价值观的表征所掩盖。结论指出,发展生成性价值-价值推理是实现可扩展机器理论心智的核心组成部分。

原文摘要 · Abstract (English)

Value learning is a crucial aspect of safe and ethical AI. This is primarily pursued by methods inferring human values from behaviour. However, humans care about much more than we are able to demonstrate through our actions. Consequently, an AI must predict the rest of our seemingly complex values from a limited sample. I call this the value generalization problem. In this paper, I argue that human values have a generative rational structure and that this allows us to solve the value generalization problem. In particular, we can use Bayesian Theory of Mind models to infer human values not only from behaviour, but also from other values. This has been obscured by the widespread use of simple utility functions to represent human values. I conclude that developing generative value-to-value inference is a crucial component of achieving a scalable machine theory of mind.

价值学习理论心智泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。