arXiv:2504.15619cs.CV2025-04被引 8

让多模态大模型更懂视觉细节,减少幻觉。

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

  • 用多视觉模型删减图像关键元素,构建更敏感的视觉偏好对。
  • 在物体幻觉基准上,幻觉率降低93.7%(响应级)和96.4%(提及级)。
  • 适合追求视觉准确性的多模态模型对齐研究者。

通过直接偏好优化(DPO)实现多模态大语言模型(MLLMs)与人类偏好的对齐已展现出显著效果。然而,现有方法主要关注语言偏好,忽视了关键的视觉上下文。本文提出自适应视觉增强偏好优化(AdaViP),通过两项关键创新解决此问题:(1) 基于视觉的偏好对构建,整合多个视觉基础模型,有策略地从图像中移除关键视觉元素,提升MLLM对视觉细节的敏感性;(2) 自适应偏好优化,动态平衡视觉与语言偏好,实现更精准的对齐。跨多个基准的广泛评估表明其有效性。值得注意的是,AdaViP-7B在Object HalBench上分别实现93.7%和96.4%的响应级与提及级幻觉减少,显著优于当前最先进方法。

原文摘要 · Abstract (English)

Preference alignment through Direct Preference Optimization (DPO) has demonstrated significant effectiveness in aligning multimodal large language models (MLLMs) with human preferences. However, existing methods focus primarily on language preferences while neglecting the critical visual context. In this paper, we propose an Adaptive Vision-enhanced Preference optimization (AdaViP) that addresses these limitations through two key innovations: (1) vision-based preference pair construction, which integrates multiple visual foundation models to strategically remove key visual elements from the image, enhancing MLLMs' sensitivity to visual details; and (2) adaptive preference optimization that dynamically balances vision- and language-based preferences for more accurate alignment. Extensive evaluations across different benchmarks demonstrate our effectiveness. Notably, our AdaViP-7B achieves 93.7% and 96.4% reductions in response-level and mentioned-level hallucination respectively on the Object HalBench, significantly outperforming current state-of-the-art methods.

多模态对齐视觉感知幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。