区分描述长度与信息密度,提升图像描述评估准确性
When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
- 用对比集定义描述特异性,分离长度与信息量
- 控制长度后,更具体描述受人青睐,无论长短
- 评估应优先特异性,而非追求冗长文字
视觉语言模型(VLMs)正越来越多地通过文本描述使视觉内容可访问。然而当前系统中,描述特异性常与长度混淆。我们认为这两者必须解耦:描述可以简洁但信息密集,也可以冗长却空洞。我们基于对比集定义特异性,即描述在多大程度上能将目标图像与其他可能图像区分开来。我们构建了一个控制长度但变化信息量的数据集,并验证了人们始终偏好更具体的描述,无论其长度如何。研究发现,仅控制长度无法解释特异性差异;长度预算的分配方式同样重要。这些结果支持直接以特异性为优先的评估方法。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used to make visual content accessible via text-based descriptions. In current systems, however, description specificity is often conflated with their length. We argue that these two concepts must be disentangled: descriptions can be concise yet dense with information, or lengthy yet vacuous. We define specificity relative to a contrast set, where a description is more specific to the extent that it picks out the target image better than other possible images. We construct a dataset that controls for length while varying information content, and validate that people reliably prefer more specific descriptions regardless of length. We find that controlling for length alone cannot account for differences in specificity: how the length budget is allocated makes a difference. These results support evaluation approaches that directly prioritize specificity over verbosity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。