用困惑度方差衡量图像描述多样性,发现评分器不同结论相反。
Surprisal reveals diversity gaps in image captioning and different scorers change the story
- 用词级困惑度方差衡量描述文本的多样性
- 人类描述的多样性是模型的两倍,但换评分器后结果反转
- 评估多样性需多评分器验证,避免单一指标误导
我们通过困惑度方差(surprisal variance)量化图像描述中的语言多样性——即同一图像生成的多个描述中,词级别负对数概率的分布范围。在MSCOCO测试集上,对比五种前沿视觉-语言大模型(贪婪与核采样生成),以及人类描述。使用基于描述训练的n-gram语言模型评估时,人类描述的困惑度方差约为模型的两倍;但若改用通用语言模型重评分,该趋势完全逆转。本研究提出一种基于困惑度的图像描述多样性评估方法,揭示单一评分器可能导致结论完全相反,因此鲁棒的多样性评估必须在多个评分器下报告结果。
原文摘要 · Abstract (English)
We quantify linguistic diversity in image captioning with surprisal variance - the spread of token-level negative log-probabilities within a caption set. On the MSCOCO test set, we compare five state-of-the-art vision-and-language LLMs, decoded with greedy and nucleus sampling, to human captions. Measured with a caption-trained n-gram LM, humans display roughly twice the surprisal variance of models, but rescoring the same captions with a general-language model reverses the pattern. Our analysis introduces the surprisal-based diversity metric for image captioning. We show that relying on a single scorer can completely invert conclusions, thus, robust diversity evaluation must report surprisal under several scorers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。