大模型能模拟人类感官感知,但机制仍不同。
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
- 用3611个词测试21个模型的感官强度判断能力
- 顶级模型相关性达0.58-0.65,准确率85%-90%
- 多模态未必带来本质提升,仍缺具身认知
本研究探讨了多模态大语言模型是否可通过统计学习实现类人感官扎根,考察其在跨感官模态中对感知强度评分的捕捉能力。通过3,611个来自Lancaster Sensorimotor Norms的词汇,评估了来自GPT、Gemini、LLaMA、Qwen四个系列共21个模型,采用相关性、距离度量与定性分析。结果表明,更大(8次比较中6次)、多模态(7次中5次)、更新的模型(8次中5次)普遍优于小规模、纯文本及旧版模型。顶尖模型达到85%-90%准确率与0.58-0.65的相关性,显示高度相似性。分布因素影响极小,未超过人类依赖水平。然而,即便表现接近,模型与人类仍存在差异,顶级模型在距离与相关性上仍有偏差,定性分析揭示其处理模式与缺失感官扎根有关。多模态虽提升性能,但似乎仅提供与海量文本相当的信息,并非质变数据,且优势出现在无关感官维度,纯文本大模型亦表现相近。研究证实,先进大模型可借统计学习逼近人类感官-语言关联,但在处理机制上仍不同于具身认知。
原文摘要 · Abstract (English)
This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics (size, multimodal capabilities, architectural generation) influence grounding performance, distributional factor dependencies (word frequency, embeddings, feature distances), and human-model processing differences. We evaluated 21 models from four families (GPT, Gemini, LLaMA, Qwen) using 3,611 words from the Lancaster Sensorimotor Norms through correlation, distance metrics, and qualitative analysis. Results showed that larger (6 out of 8 comparisons), multimodal (5 of 7), and newer models (5 of 8) generally outperformed their smaller, text-based, and older counterparts. Top models achieved 85-90% accuracy and 0.58-0.65 correlations with human ratings, demonstrating substantial similarity. Moreover, distributional factors showed minimal impact, not exceeding human dependency levels. However, despite strong alignment, models were not identical to humans, as even top performers showed differences in distance and correlation measures, with qualitative analysis revealing processing patterns related to absent sensory grounding. Additionally, it remains questionable whether introducing multimodality resolves this grounding deficit. Although multimodality improved performance, it seems to provide similar information as massive text rather than qualitatively different data, as benefits occurred across unrelated sensory dimensions and massive text-only models achieved comparable results. Our findings demonstrate that while advanced LLMs can approximate human sensory-linguistic associations through statistical learning, they still differ from human embodied cognition in processing mechanisms, even with multimodal integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。