arXiv:2506.11631cs.CL2025-06ACL被引 3

人类对拼图形状的描述受场景影响,模型却忽视这种多样性。

SceneGram: Conceptualizing and Describing Tangrams in Scene Context

  • 构建场景拼图数据集,研究场景如何影响人们对抽象形状的命名
  • 发现人类对同一拼图可有多种概念化命名,如'螃蟹'或'飞船'
  • 对比大模型生成结果,揭示其缺乏人类概念的丰富性与情境依赖

认知科学研究表明,人类对同一物体可能有多种不同的概念化和命名方式,例如同一个抽象拼图形状可被称作‘螃蟹’、‘水槽’或‘宇宙飞船’。另一个普遍假设是,场景上下文会深刻影响我们的视觉感知和概念预期。本文提出SceneGram数据集,收集了人类在不同场景中对拼图形状的参考描述,支持对场景上下文如何影响概念化的系统分析。基于该数据,我们评估了多模态大语言模型生成的参考描述,发现这些模型未能捕捉到人类参考中所呈现的概念丰富性和变异性。

原文摘要 · Abstract (English)

Research on reference and naming suggests that humans can come up with very different ways of conceptualizing and referring to the same object, e.g. the same abstract tangram shape can be a "crab", "sink" or "space ship". Another common assumption in cognitive science is that scene context fundamentally shapes our visual perception of objects and conceptual expectations. This paper contributes SceneGram, a dataset of human references to tangram shapes placed in different scene contexts, allowing for systematic analyses of the effect of scene context on conceptualization. Based on this data, we analyze references to tangram shapes generated by multimodal LLMs, showing that these models do not account for the richness and variability of conceptualizations found in human references.

概念化拼图多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。