arXiv:2412.15739cs.CV2024-12被引 3

通过图像序关系校准,减少大模型幻觉生成错误物体。

VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models

  • 基于修改图像对的序关系校准令牌预测
  • 在多个基准上显著降低物体幻觉率
  • 无需训练的轻量级方法,适合部署优化

大型视觉语言模型(LVLM)在大语言模型兴起后取得显著进展。然而,这些模型常基于源内容生成看似合理但不准确或不一致的信息,即“幻觉”,可能在实际应用中造成严重后果。为此,本文提出VORD,一种简单有效的校准方法,通过分析修改图像对间的序关系来缓解幻觉。VORD有两种形式:1)无需训练的极简变体,可剔除修改图像对中不合理的令牌;2)可训练的目标函数,对不合理令牌施加惩罚。实验表明,VORD在多种LVLM基准上均实现更好校准效果,有效抑制物体幻觉。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have made remarkable developments along with the recent surge of large language models. Despite their advancements, LVLMs have a tendency to generate plausible yet inaccurate or inconsistent information based on the provided source content. This phenomenon, also known as ``hallucinations" can have serious downstream implications during the deployment of LVLMs. To address this, we present VORD a simple and effective method that alleviates hallucinations by calibrating token predictions based on ordinal relationships between modified image pairs. VORD is presented in two forms: 1.) a minimalist training-free variant which eliminates implausible tokens from modified image pairs, and 2.) a trainable objective function that penalizes unlikely tokens. Our experiments demonstrate that VORD delivers better calibration and effectively mitigates object hallucinations on a wide-range of LVLM benchmarks.

视觉语言模型幻觉抑制校准方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。