arXiv:2507.19052cs.CV2025-07

用视听特征建模大脑对自然多模态刺激的响应,发现简单模型更泛化。

Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding

  • 用X-CLIP和Whisper提取视听特征,构建脑编码模型。
  • 线性模型在新场景下比注意力模型高18%准确率。
  • 语言特征无效,说明大脑更依赖连续视听流。

预测大脑对自然多模态刺激的反应是计算神经科学的关键挑战。本文采用先进的视觉(X-CLIP)和听觉(Whisper)特征提取器构建脑编码模型,并在分布内(ID)和多样分布外(OOD)数据上严格评估。结果揭示了模型复杂度与泛化能力间的根本权衡:高容量注意力模型在ID数据上表现优异,但简单线性模型更具鲁棒性,在OOD集上优于竞争基线18%。令人意外的是,语言特征未提升预测精度,表明对于熟悉语言,神经编码可能由连续视听流主导,而非冗余文本信息。空间分析显示,模型在听觉皮层表现显著提升,凸显高质量语音表征的优势。研究证明,严格的OOD测试对构建稳健神经-AI模型至关重要,并深化了对模型架构、刺激特性与感官层次如何塑造多模态世界神经编码的理解。

原文摘要 · Abstract (English)

Predicting brain activity in response to naturalistic, multimodal stimuli is a key challenge in computational neuroscience. While encoding models are becoming more powerful, their ability to generalize to truly novel contexts remains a critical, often untested, question. In this work, we developed brain encoding models using state-of-the-art visual (X-CLIP) and auditory (Whisper) feature extractors and rigorously evaluated them on both in-distribution (ID) and diverse out-of-distribution (OOD) data. Our results reveal a fundamental trade-off between model complexity and generalization: a higher-capacity attention-based model excelled on ID data, but a simpler linear model was more robust, outperforming a competitive baseline by 18\% on the OOD set. Intriguingly, we found that linguistic features did not improve predictive accuracy, suggesting that for familiar languages, neural encoding may be dominated by the continuous visual and auditory streams over redundant textual information. Spatially, our approach showed marked performance gains in the auditory cortex, underscoring the benefit of high-fidelity speech representations. Collectively, our findings demonstrate that rigorous OOD testing is essential for building robust neuro-AI models and provides nuanced insights into how model architecture, stimulus characteristics, and sensory hierarchies shape the neural encoding of our rich, multimodal world.

脑科学多模态神经编码泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。