arXiv:2511.15895cs.AI2025-11被引 1

发现大模型的共情能力比逻辑推理更关键,能提升其理解他人想法的能力。

Decomposing Theory of Mind: How Emotional Processing Mediates ToM Abilities in LLMs

  • 通过对比调节前后模型激活,拆解共情在大模型理解他人信念中的作用
  • 调节后信念判断准确率从32.5%升至46.7%,主要依赖情绪感知与价值判断
  • 适合关注模型心理建模、认知机制解释的研究者阅读

近期研究显示,激活调节可显著提升语言模型的理论心智(ToM)能力(Bortoletto et al. 2024),但其内部机制尚不明确。本文通过线性探测45种认知行为,比较调节前后模型激活差异,采用对比激活添加(CAA)调节Gemma-3-4B,在1,000个BigToM前向信念场景上评估。结果发现,信念归因任务准确率从32.5%提升至46.7%,该提升由情绪处理机制驱动:情绪感知得分+2.23,情绪价值判断+2.20;而分析性过程减弱:质疑行为-0.78,收敛思维-1.59。表明大模型的理论心智能力主要由情绪理解而非逻辑推理中介。

原文摘要 · Abstract (English)

Recent work shows activation steering substantially improves language models' Theory of Mind (ToM) (Bortoletto et al. 2024), yet the mechanisms of what changes occur internally that leads to different outputs remains unclear. We propose decomposing ToM in LLMs by comparing steered versus baseline LLMs' activations using linear probes trained on 45 cognitive actions. We applied Contrastive Activation Addition (CAA) steering to Gemma-3-4B and evaluated it on 1,000 BigToM forward belief scenarios (Gandhi et al. 2023), we find improved performance on belief attribution tasks (32.5\% to 46.7\% accuracy) is mediated by activations processing emotional content : emotion perception (+2.23), emotion valuing (+2.20), while suppressing analytical processes: questioning (-0.78), convergent thinking (-1.59). This suggests that successful ToM abilities in LLMs are mediated by emotional understanding, not analytical reasoning.

理论心智情绪理解大模型机制认知解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。