用脑扫描信号直接生成带情绪的描述文本,更贴近真实感受。
EmoMind: Decoding Affective Captions from Human Brain fMRI

- 从fMRI信号中解码连续情绪向量,重写视觉描述。
- 在两个数据集上优于用离散标签提示的GPT-4,尤其在个体化情绪表达上。
- 支持个体情绪结构建模,适合研究个性化情感机制的人。
从脑活动解码视觉体验已取得显著进展,但现有脑到文本系统多恢复语义内容而忽略情感。语言模型虽能通过类别标签生成情绪文本,但此类标签将丰富的个体差异压缩为粗粒度离散类别。我们提出EmoMind,首个端到端从fMRI信号解码情感化描述的框架。该方法先从脑激活中还原语义中性场景描述,再利用同一fMRI记录解码的34维连续情绪向量重写描述。通过无分类器引导训练,以保持语义一致性的同时调节情感表达强度,实现语义保真与情感丰富性的平滑过渡。我们在三轴验证框架下评估:个体特异性、结构几何性和因果控制性。进一步引入合成脑替换测试以检验对测量设备的鲁棒性,并以提示脑解码前5个情绪标签的GPT-4作为强基线。在两个独立的情绪fMRI数据集上,EmoMind在所有指标上均显著优于基线,尤其在需要个体化情绪结构的指标上提升最大。结果表明,连续脑解码情绪可作为个性化情感描述的有效控制信号,为研究个体情感脑组织开辟新路径。
原文摘要 · Abstract (English)
Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic content while discarding affect. Additionally, language models can generate emotional text when prompted with categorical labels, but such labels collapse rich inter-subject variability into coarse discrete bins. We present EmoMind, the first end-to-end pipeline for decoding affective captions directly from fMRI signals. EmoMind first retrieves a semantically grounded neutral scene description from brain-decoded visual features, then rewrites it using a continuous 34-dimensional emotion vector decoded from the same fMRI recording. To control the balance between content preservation and affective expression, we train the rewriter with classifier-free guidance against an identity-preserving null branch, enabling smooth interpolation between semantic fidelity and affective expressivity. We evaluate affective caption generation with a three-axis validation framework spanning subject-specificity, structural geometry, and causal control. We further augment this framework with a synthetic-brain substitution test that probes robustness to the measurement apparatus, and we benchmark each axis against GPT-4 prompted with brain-decoded top-5 emotion labels as a strong discrete baseline. Across two independent emotion fMRI datasets, EmoMind significantly outperforms label-prompted GPT-4 on all three axes, with the largest gains on metrics that require person-specific affective structure rather than population-level emotion aggregation. These results establish continuous brain-decoded affect as a viable control signal for individualized affective caption generation and open new directions for studying individual affective brain organisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。