对比大模型与人类在数量表达理解上的差异,揭示模型认知短板。
Quantification and object perception in Multimodal Large Language Models and human linguistic cognition
- 从人类语言共性出发,分析数量词的层级、使用范围和心理数感偏差。
- 模型在数量估计和量词排序上表现良好,但使用范围和典型性判断仍有差距。
- 跨语言比较揭示模型能力是否稳定,适合认知科学与AI交叉研究者。
数量表达是(多模态)大语言模型(MLLMs)的一项挑战性语言现象。尽管其涉及逻辑、语用和数值领域,但性能不佳的具体原因仍不明确。本文聚焦人类语言中普遍存在的三个未被充分研究的数量特征:量词的层级排列、使用范围与典型性,以及人类近似数系统中的固有偏差。旨在探究这些特征如何被模型架构编码,与人类有何差异,并评估模型类型(思考型 vs. 指令型)及语言对结果的影响。结果显示,思考型模型在数量估计和量词层级组织任务中表现优异,但所有模型类型在使用范围和典型性判断上仍存在显著人类-模型差异。该研究为理解多模态大模型作为语义与语用代理的本质提供了新路径,跨语言视角亦有助于检验其能力的鲁棒性与稳定性。
原文摘要 · Abstract (English)
Quantification has been proven to be a particularly difficult linguistic phenomenon for (Multimodal) Large Language Models (MLLMs). However, given that quantification interfaces with the logic, pragmatic, and numerical domains, the exact reasons for the poor performance are still unclear. This paper looks at three key features of human quantification shared cross-linguistically that have remained so far unexplored in the (M)LLM literature: the ordering of quantifiers into scales, the ranges of use and prototypicality, and the biases inherent in the human approximate number system. The aim is to determine how these features are encoded in the models' architecture, how they may differ from humans, and whether the results are affected by the type of model (thinking vs. instruct) and the language under investigation. Results show that although thinking models showed a high accuracy in the numerosity estimation task and in the organization of quantifiers into scales, there are still key differences between humans and LLMs across all model types, particularly in terms of ranges of use and prototypicality values. This work, thus, paves the way for addressing the nature of MLLMs as semantic and pragmatic agents, while the cross-linguistic lens can elucidate whether their abilities are robust and stable across different languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。