通过自校准视觉锚定奖励,实现细粒度防幻觉优化。
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
- 基于图像与噪声图的生成词元分布差,构建可自校准的词元级奖励
- 在LLaVA-1.5-7B上,幻觉评测指标提升显著,绝对增益明显
- 无需人工标注,自动聚焦视觉相关词元,适合多模态模型优化
直接偏好优化(DPO)已被证明能有效缓解大视觉语言模型(LVLMs)中的幻觉问题,通过使模型输出更贴近人类偏好。然而现有方法存在两大缺陷:1)缺乏可扩展的词元级奖励;2)忽略视觉锚定词元。为此,本文提出一种新型词元偏好优化模型TPO,通过自校准奖励机制,无需细粒度标注即可自适应关注与视觉相关的词元。具体地,引入词元级的视觉锚定奖励,计算在原始图像和扰动图像条件下生成词元的逻辑分布差异。此外,设计视觉感知训练目标以增强对关键视觉锚定词元的识别,实现更精确的词元级优化。大量实验表明,所提TPO达到当前最优性能。例如,在基于LLaVA-1.5-7B的基础上,其在幻觉基准测试中取得显著的绝对性能提升。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has been demonstrated to be highly effective in mitigating hallucinations in Large Vision Language Models (LVLMs) by aligning their outputs more closely with human preferences. Despite the recent progress, existing methods suffer from two drawbacks: 1) Lack of scalable token-level rewards; and 2) Neglect of visual-anchored tokens. To this end, we propose a novel Token Preference Optimization model with self-calibrated rewards (dubbed as TPO), which adaptively attends to visual-correlated tokens without fine-grained annotations. Specifically, we introduce a token-level \emph{visual-anchored} \emph{reward} as the difference of the logistic distributions of generated tokens conditioned on the raw image and the corrupted one. In addition, to highlight the informative visual-anchored tokens, a visual-aware training objective is proposed to enhance more accurate token-level optimization. Extensive experimental results have manifested the state-of-the-art performance of the proposed TPO. For example, by building on top of LLAVA-1.5-7B, our TPO boosts the performance absolute improvement for hallucination benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。