arXiv:2505.17024cs.AIq-bio.NC2025-05被引 1

用情绪趋性理论解释智能体对齐问题,为理解价值观形成提供新视角。

An Affective-Taxis Hypothesis for Alignment and Interpretability

  • 将目标与价值重新定义为情绪趋性,结合神经科学构建计算模型。
  • 在可简化生物模型中验证了该理论对趋性导航的解释力。
  • 适合关注AI对齐、认知机制与情感建模的研究者阅读。

人工智能对齐研究旨在开发方法,确保智能体无论能力如何,始终以符合其人类操作者目标与价值观的方式行动。本文提出一种情感主义对齐思路,将目标与价值观重新表述为情绪趋性,并通过进化发育与计算神经科学的最新成果解释情绪效价的产生机制。我们综述了该领域的前沿进展,并在此基础上提出一个基于趋性导航的计算情感模型。该模型在可简化模型生物中得到了证据支持,显示出对生物趋性导航的合理反映。最后,讨论了情绪趋性在人工智能对齐中的潜在作用。

原文摘要 · Abstract (English)

AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This paper proposes an affectivist approach to the alignment problem, re-framing the concepts of goals and values in terms of affective taxis, and explaining the emergence of affective valence by appealing to recent work in evolutionary-developmental and computational neuroscience. We review the state of the art and, building on this work, we propose a computational model of affect based on taxis navigation. We discuss evidence in a tractable model organism that our model reflects aspects of biological taxis navigation. We conclude with a discussion of the role of affective taxis in AI alignment.

AI对齐情绪建模认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。