视觉输入未必提升大模型对身体知识的理解。
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
- 构建多感官感知基准,测试模型在视觉等五感上的理解能力。
- 30个主流模型中,视觉语言模型在各任务上均不如纯文本模型。
- 模型对空间推理和触觉感知表现差,向量表示易受词频干扰。
尽管多模态语言模型取得显著进展,但视觉接地是否增强其对身体知识的理解仍不明确。为此,我们基于心理学中的感知理论,提出一个全新的身体知识理解基准,涵盖视觉、听觉、触觉、味觉、嗅觉及内感受等外部感官与内部感知。该基准通过向量对比和问答任务(超1700个问题)评估模型在不同感官维度的感知能力。对比30个前沿语言模型发现,视觉语言模型在两项任务中均未优于纯文本模型,且在视觉维度表现显著更差。进一步分析表明,向量表示易受词汇形式与频率影响,模型在空间感知与推理任务上表现薄弱。研究强调需更有效融合身体知识以提升模型对物理世界的理解。
原文摘要 · Abstract (English)
Despite significant progress in multimodal language models (LMs), it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models. To address this question, we propose a novel embodied knowledge understanding benchmark based on the perceptual theory from psychology, encompassing visual, auditory, tactile, gustatory, olfactory external senses, and interoception. The benchmark assesses the models' perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. By comparing 30 state-of-the-art LMs, we surprisingly find that vision-language models (VLMs) do not outperform text-only models in either task. Moreover, the models perform significantly worse in the visual dimension compared to other sensory dimensions. Further analysis reveals that the vector representations are easily influenced by word form and frequency, and the models struggle to answer questions involving spatial perception and reasoning. Our findings underscore the need for more effective integration of embodied knowledge in LMs to enhance their understanding of the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。