用类脑共情机制让AI自发利他,提升道德决策稳定性。
Building Altruistic and Moral AI Agent with Brain-inspired Emotional Empathy Mechanisms
- 模拟人类共情神经回路,通过情绪共鸣驱动内在利他动机。
- 在三类场景中稳定表现利他行为,共情水平与利他偏好正相关。
- 适用于伦理困境、对抗环境等复杂场景,适合安全型AI研发者。
随着AI日益融入人类社会,确保其行为安全、利他并符合人类道德价值至关重要。现有研究在将伦理嵌入AI方面仍显不足,基于规则的外部约束难以提供长期稳定性和泛化能力。情绪共情通过情绪共享与传播机制,天然激发利他行为。受此启发,本文借鉴人类共情驱动利他决策的神经机制,模拟自我-他人感知-镜像-共情神经回路,构建类脑共情驱动的利他决策模型。该模型中,共情直接调控多巴胺释放,形成内在利他动机。实验显示,模型在三类设置下均表现出一致利他行为:融合情绪传染的双智能体救援、多智能体博弈及机器人情感互动。深入分析验证了共情水平与利他偏好间的正相关关系(符合心理学实验结果),并揭示交互对象共情水平对智能体行为模式的影响。进一步测试表明,该模型在自利与他人福祉冲突、部分可观测环境及对抗防御场景中具备良好性能与稳定性。本工作为实现类人共情驱动的道德决策提供了初步探索,为发展伦理对齐的AI提供新视角。
原文摘要 · Abstract (English)
As AI closely interacts with human society, it is crucial to ensure that its behavior is safe, altruistic, and aligned with human ethical and moral values. However, existing research on embedding ethical considerations into AI remains insufficient, and previous external constraints based on principles and rules are inadequate to provide AI with long-term stability and generalization capabilities. Emotional empathy intrinsically motivates altruistic behaviors aimed at alleviating others' negative emotions through emotional sharing and contagion mechanisms. Motivated by this, we draw inspiration from the neural mechanism of human emotional empathy-driven altruistic decision making, and simulate the shared self-other perception-mirroring-empathy neural circuits, to construct a brain-inspired emotional empathy-driven altruistic decision-making model. Here, empathy directly impacts dopamine release to form intrinsic altruistic motivation. The proposed model exhibits consistent altruistic behaviors across three experimental settings: emotional contagion-integrated two-agent altruistic rescue, multi-agent gaming, and robotic emotional empathy interaction scenarios. In-depth analyses validate the positive correlation between empathy levels and altruistic preferences (consistent with psychological behavioral experiment findings), while also demonstrating how interaction partners' empathy levels influence the agent's behavioral patterns. We further test the proposed model's performance and stability in moral dilemmas involving conflicts between self-interest and others' well-being, partially observable environments, and adversarial defense scenarios. This work provides preliminary exploration of human-like empathy-driven altruistic moral decision making, contributing potential perspectives for developing ethically-aligned AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。