用问答方式让视觉语言模型为维基图片打分,提升图文匹配度。
QuizRank: Picking Images by Quizzing VLMs
- 将文章内容转为多选题,用VLM答题表现评估图片质量。
- 在12个主题上人类评分与模型打分相关性达0.87,显著优于基线。
- 适合需要精准图文匹配的科普内容创作与编辑辅助场景。
图像在提升维基百科文章可读性和理解度方面发挥关键作用,但并非所有图像效果相同,且并非所有编辑都具备图片选择能力。我们提出QuizRank,一种利用大语言模型(LLMs)和视觉语言模型(VLMs)对图像进行排序的新方法,将其作为学习干预手段。该方法将文章主题的文本描述转化为关于概念重要视觉特征的多项选择题,通过让VLM回答这些问题来评估图像表现:能更好帮助回答问题的图像获得更高排名。为进一步增强对视觉相似图像的区分能力,我们引入对比式QuizRank,利用目标概念(如西蓝鸟)与干扰项概念(如山蓝鸟)之间的特征差异生成问题。实验表明,VLM在人类答题者评分中表现出高度一致性(相关系数0.87),并能有效区分图像质量,验证其作为高效视觉评价工具的潜力。
原文摘要 · Abstract (English)
Images play a vital role in improving the readability and comprehension of Wikipedia articles by serving as `illustrative aids.' However, not all images are equally effective and not all Wikipedia editors are trained in their selection. We propose QuizRank, a novel method of image selection that leverages large language models (LLMs) and vision language models (VLMs) to rank images as learning interventions. Our approach transforms textual descriptions of the article's subject into multiple-choice questions about important visual characteristics of the concept. We utilize these questions to quiz the VLM: the better an image can help answer questions, the higher it is ranked. To further improve discrimination between visually similar items, we introduce a Contrastive QuizRank that leverages differences in the features of target (e.g., a Western Bluebird) and distractor concepts (e.g., Mountain Bluebird) to generate questions. We demonstrate the potential of VLMs as effective visual evaluators by showing a high congruence with human quiz-takers and an effective discriminative ranking of images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。