大模型如何理解物体的感官属性?研究发现视觉与语言模型各有优势。
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
- 用探测任务测试大模型对物体属性的感知能力,涵盖颜色、气味等多维度
- 多模态模型略优于纯语言模型,但纯视觉模型在非视觉属性上表现不俗
- 揭示了单模态学习的潜力与多模态互补性,适合关注模型认知机制的研究者
人类学习和概念表征基于感官运动体验,而当前主流基础模型则不然。本文探讨大规模模型在海量数据训练下,对具体物体概念(如玫瑰是红色、有甜香、属于花类)的语义特征规范的表征能力。通过探测任务,评估仅使用图像数据训练的图像编码器、多模态训练的图像编码器以及纯语言模型,在预测扩展版经典McRae属性规范和最新Binder属性数据集上的表现。结果表明,多模态图像编码器略优于纯语言模型,而仅图像训练的编码器在非视觉属性(如'百科知识'或'功能'类)上表现与语言模型相当。这些发现为单模态学习的潜力及模态间的互补性提供了新见解。
原文摘要 · Abstract (English)
Human learning and conceptual representation is grounded in sensorimotor experience, in contrast to state-of-the-art foundation models. In this paper, we investigate how well such large-scale models, trained on vast quantities of data, represent the semantic feature norms of concrete object concepts, e.g. a ROSE is red, smells sweet, and is a flower. More specifically, we use probing tasks to test which properties of objects these models are aware of. We evaluate image encoders trained on image data alone, as well as multimodally-trained image encoders and language-only models, on predicting an extended denser version of the classic McRae norms and the newer Binder dataset of attribute ratings. We find that multimodal image encoders slightly outperform language-only approaches, and that image-only encoders perform comparably to the language models, even on non-visual attributes that are classified as "encyclopedic" or "function". These results offer new insights into what can be learned from pure unimodal learning, and the complementarity of the modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。