发现人工神经网络可自发形成类镜像神经元模式,助力AI内生式对齐。
Mirror-Neuron Patterns in AI Alignment
- 设计蛙与蟾蜍游戏框架,诱导模型发展共享动作表征。
- 在适当规模与耦合条件下,模型激活强度达0.62(CMNI)并支持合作行为。
- 适合关注内生对齐、社会智能与可解释性的研究者参考。
随着人工智能逼近超人能力,使其与人类价值观对齐变得愈发关键。现有对齐策略依赖外部约束,面对未来超智能系统可能失效。本研究探究人工神经网络是否能发展出类似生物镜像神经元的模式——即执行与观察动作时均激活,并分析其对内在对齐的潜在贡献。镜像神经元在人类共情、模仿与社会认知中起核心作用。研究构建了促进合作行为的‘蛙与蟾蜍’游戏框架,识别镜像神经元模式出现的条件,评估其对动作电路的影响,提出检查点镜像神经元指数(CMNI)量化激活强度与一致性,并建立理论框架。结果表明,恰当规模的模型容量与自我/他者耦合可促成类似镜像神经元的共享神经表征。这些类共情回路支持合作行为,提示通过镜像神经动力学建模内在动机,或可将类共情机制直接嵌入AI架构,补充现有对齐技术。
原文摘要 · Abstract (English)
As artificial intelligence (AI) advances toward superhuman capabilities, aligning these systems with human values becomes increasingly critical. Current alignment strategies rely largely on externally specified constraints that may prove insufficient against future super-intelligent AI capable of circumventing top-down controls. This research investigates whether artificial neural networks (ANNs) can develop patterns analogous to biological mirror neurons cells that activate both when performing and observing actions, and how such patterns might contribute to intrinsic alignment in AI. Mirror neurons play a crucial role in empathy, imitation, and social cognition in humans. The study therefore asks: (1) Can simple ANNs develop mirror-neuron patterns? and (2) How might these patterns contribute to ethical and cooperative decision-making in AI systems? Using a novel Frog and Toad game framework designed to promote cooperative behaviors, we identify conditions under which mirror-neuron patterns emerge, evaluate their influence on action circuits, introduce the Checkpoint Mirror Neuron Index (CMNI) to quantify activation strength and consistency, and propose a theoretical framework for further study. Our findings indicate that appropriately scaled model capacities and self/other coupling foster shared neural representations in ANNs similar to biological mirror neurons. These empathy-like circuits support cooperative behavior and suggest that intrinsic motivations modeled through mirror-neuron dynamics could complement existing alignment techniques by embedding empathy-like mechanisms directly within AI architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。