arXiv:2502.11874cs.CL2025-02ACL被引 4

研究视觉模型如何理解模糊数量词,发现其表现与人类相似但存在不一致。

VAQUUM: Are Vague Quantifiers Grounded in Visual Data?

  • 构建新数据集VAQUUM,含20,300条人类对图像中数量描述的评分
  • 模型判断模糊数量词时受物体数量影响,与人类模式一致
  • 不同评估方式下模型表现差异大,说明生成与判断是两种不同机制

模糊数量词如'少量'和'许多'受上下文因素影响,包括场景中物体的数量。本文评估视觉语言模型(VLMs)在视觉语境中生成或判断模糊数量词时与人类的一致性。我们发布了新数据集VAQUUM,包含1089张图像上20,300条人类对数量陈述的评分。通过三种评估方法比较人类判断与VLM预测。结果表明,VLMs与人类一样受物体数量影响,但在不同评估设置中表现出显著不一致性,提示判断与生成模糊数量词依赖于两种不同过程。

原文摘要 · Abstract (English)

Vague quantifiers such as "a few" and "many" are influenced by various contextual factors, including the number of objects present in a given context. In this work, we evaluate the extent to which vision-and-language models (VLMs) are compatible with humans when producing or judging the appropriateness of vague quantifiers in visual contexts. We release a novel dataset, VAQUUM, containing 20,300 human ratings on quantified statements across a total of 1089 images. Using this dataset, we compare human judgments and VLM predictions using three different evaluation methods. Our findings show that VLMs, like humans, are influenced by object counts in vague quantifier use. However, we find significant inconsistencies across models in different evaluation settings, suggesting that judging and producing vague quantifiers rely on two different processes.

视觉语言模型模糊量化人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。