语言模型比多模态模型更贴近人类真实体验和脑活动模式。
Experiential Semantic Information and Brain Alignment: Are Multimodal Models Better than Language Models?
- 用经验性语义模型对比文本表示的现实感知能力
- 语言模型在捕捉体验信息和脑响应对齐上更优
- 适合关注模型与人脑关联性的研究者
计算语言学普遍认为,基于图像或音频的多模态模型所学文本表征比纯语言模型更丰富、更接近人类认知,因其具备真实世界体验的根基。然而,验证这一假设的实证研究仍很缺乏。本文通过对比对比学习型多模态模型与纯语言模型的词表示,在捕捉经验性语义信息(基于已有范式化的‘经验模型’)以及与人类fMRI脑响应对齐程度上的表现,发现令人意外的结果:语言模型在这两方面均优于多模态模型。此外,语言模型还学习到更多与大脑相关但未被经验模型覆盖的语义信息。本研究揭示了当前多模态模型整合跨模态信息的不足,强调需发展更有效融合多源语义信息的计算模型。
原文摘要 · Abstract (English)
A common assumption in Computational Linguistics is that text representations learnt by multimodal models are richer and more human-like than those by language-only models, as they are grounded in images or audio -- similar to how human language is grounded in real-world experiences. However, empirical studies checking whether this is true are largely lacking. We address this gap by comparing word representations from contrastive multimodal models vs. language-only ones in the extent to which they capture experiential information -- as defined by an existing norm-based 'experiential model' -- and align with human fMRI responses. Our results indicate that, surprisingly, language-only models are superior to multimodal ones in both respects. Additionally, they learn more unique brain-relevant semantic information beyond that shared with the experiential model. Overall, our study highlights the need to develop computational models that better integrate the complementary semantic information provided by multimodal data sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。