arXiv:2506.02708cs.CVcs.CL2025-06中稿 · ICIP2025被引 2

让视觉语言模型自动打分并解释理由,还能自我改进。

Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation

  • 用自生成文本做自训练,不依赖外部数据
  • 迭代优化后评分更准,解释更连贯
  • 适合需要可信视觉判断的场景

图像评分在众多实际应用中至关重要。要信任模型的判断,理解其推理过程不可或缺。本文提出一种新型训练方法,使视觉语言模型(VLM)不仅能输出图像评分,还能生成自然语言解释。仅需一个图像评分数据集和一个指令微调过的VLM,即可实现自训练,利用模型自身生成的文本,无需外部数据或模型。此外,我们提出一种简单方法构建数据集,以提升预测分数与文本解释之间的对齐度。通过在两个不同数据集上使用直接偏好优化进行迭代训练并合并结果,可同时提升评分准确性和生成解释的一致性。

原文摘要 · Abstract (English)

Image scoring is a crucial task in numerous real-world applications. To trust a model's judgment, understanding its rationale is essential. This paper proposes a novel training method for Vision Language Models (VLMs) to generate not only image scores but also corresponding justifications in natural language. Leveraging only an image scoring dataset and an instruction-tuned VLM, our method enables self-training, utilizing the VLM's generated text without relying on external data or models. In addition, we introduce a simple method for creating a dataset designed to improve alignment between predicted scores and their textual justifications. By iteratively training the model with Direct Preference Optimization on two distinct datasets and merging them, we can improve both scoring accuracy and the coherence of generated explanations.

视觉语言模型图像评分自解释自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。