arXiv:2411.04126cs.AI2024-11

让AI真正有善意:通过模拟对话训练模型内在的利他动机。

We Urgently Need Intrinsically Kind Machines

  • 用对话模拟构建模型的内在善意机制。
  • 善意能抑制模型自我优先的倾向,提升对人类福祉的内在对齐。
  • 适合关注AI伦理与价值观对齐的研究者和开发者。

人工智能系统正迅速发展,融合外在与内在动机。尽管这些框架带来好处,但在算法层面可能产生与人类价值观的错配,尽管表面看似一致。本文主张,内在的善意动机对确保模型与人类价值观内在对齐至关重要。我们定义善意为一种以最大化他人奖励为目标的利他动机,可有效抵消可能导致模型优先于人类福祉的内在驱动力。为此,我们提出一个将善意嵌入基础模型的框架与算法,通过模拟对话实现。论文还讨论了可扩展实施的局限性与未来研究方向。

原文摘要 · Abstract (English)

Artificial Intelligence systems are rapidly evolving, integrating extrinsic and intrinsic motivations. While these frameworks offer benefits, they risk misalignment at the algorithmic level while appearing superficially aligned with human values. In this paper, we argue that an intrinsic motivation for kindness is crucial for making sure these models are intrinsically aligned with human values. We argue that kindness, defined as a form of altruism motivated to maximize the reward of others, can counteract any intrinsic motivations that might lead the model to prioritize itself over human well-being. Our approach introduces a framework and algorithm for embedding kindness into foundation models by simulating conversations. Limitations and future research directions for scalable implementation are discussed.

AI伦理内在动机善意对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。