arXiv:2509.24298cs.HCcs.AI2025-09

用多模态AI模拟人类情绪,比自报更准确预测大脑活动。

Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports

  • 用AI对2180段视频做三元组判断,生成30维情绪嵌入。
  • 多模态AI的情绪表示能最准确预测人脑情绪区的神经活动。
  • 视觉信息让AI情绪模型更贴近真实大脑机制,适合研究情感神经基础。

情绪表征在人类认知与社会互动中至关重要,但其高维几何结构及其神经基础仍存争议。核心挑战在于‘行为-神经鸿沟’:人类自报难以预测脑活动。本文假设该鸿沟源于传统评分量表的局限性,并提出大规模相似性判断可更忠实捕捉情绪的脑几何结构。我们使用人工智能模型作为‘认知代理’,从多模态大语言模型(MLLM)和仅语言模型(LLM)处收集了数百万条三元组奇数项判断,针对2,180个情绪诱发视频。结果发现,这些模型生成的30维嵌入高度可解释,主要沿类别化维度组织,但融合了维度特征。最显著的是,MLLM的情绪表示在预测人类情绪加工网络神经活动方面表现最优,不仅优于LLM,甚至超越了直接来自人类行为评分的表示。这一结果支持核心假设,表明感官接地——从丰富视觉数据中学习——对于构建真正神经对齐的情绪概念框架至关重要。研究证明,多模态大模型可自主发展出丰富且神经对齐的情绪表征,为弥合主观体验与神经基质间的差距提供了强大范式。

原文摘要 · Abstract (English)

The ability to represent emotion plays a significant role in human cognition and social interaction, yet the high-dimensional geometry of this affective space and its neural underpinnings remain debated. A key challenge, the `behavior-neural gap,' is the limited ability of human self-reports to predict brain activity. Here we test the hypothesis that this gap arises from the constraints of traditional rating scales and that large-scale similarity judgments can more faithfully capture the brain's affective geometry. Using AI models as `cognitive agents,' we collected millions of triplet odd-one-out judgments from a multimodal large language model (MLLM) and a language-only model (LLM) in response to 2,180 emotionally evocative videos. We found that the emergent 30-dimensional embeddings from these models are highly interpretable and organize emotion primarily along categorical lines, yet in a blended fashion that incorporates dimensional properties. Most remarkably, the MLLM's representation predicted neural activity in human emotion-processing networks with the highest accuracy, outperforming not only the LLM but also, counterintuitively, representations derived directly from human behavioral ratings. This result supports our primary hypothesis and suggests that sensory grounding--learning from rich visual data--is critical for developing a truly neurally-aligned conceptual framework for emotion. Our findings provide compelling evidence that MLLMs can autonomously develop rich, neurally-aligned affective representations, offering a powerful paradigm to bridge the gap between subjective experience and its neural substrates. Project page: https://reedonepeck.github.io/ai-emotion.github.io/.

情绪建模多模态AI神经科学大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。