提出CATCH方法,用对比解码减少视觉语言模型的幻觉问题。
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
- 基于信息瓶颈理论,分步分离视觉信息、检测幻觉、自适应修正。
- 在多个视觉问答任务上显著降低幻觉率,无需额外训练即可泛化。
- 适合医疗、自动驾驶等对准确率要求高的关键场景使用。
大型视觉语言模型(LVLM)虽具备出色的视觉语言推理能力,但在医疗和自动驾驶等关键领域仍存在普遍且严重的幻觉问题。现有方法未能解决视觉-语言错位导致的视觉缺陷,限制了细粒度特征感知。为此,本文提出一种名为CATCH的互补自适应令牌级对比解码方法,其核心包括:互补视觉解耦(CVD)用于分离视觉信息,非视觉筛选(NVS)用于检测幻觉,以及自适应令牌级对比解码(ATCD)以抑制幻觉。CATCH有效缓解了因视觉缺陷导致的细粒度感知下降与开放场景中的累积幻觉问题。该方法适用于多种视觉问答任务,无需特定数据或先验知识,且无需额外训练即可良好泛化至新任务,为提升LVLM在复杂应用中的可靠性开辟新路径。
原文摘要 · Abstract (English)
Large Vision-Language Model (LVLM) systems have demonstrated impressive vision-language reasoning capabilities but suffer from pervasive and severe hallucination issues, posing significant risks in critical domains such as healthcare and autonomous systems. Despite previous efforts to mitigate hallucinations, a persistent issue remains: visual defect from vision-language misalignment, creating a bottleneck in visual processing capacity. To address this challenge, we develop Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs (CATCH), based on the Information Bottleneck theory. CATCH introduces Complementary Visual Decoupling (CVD) for visual information separation, Non-Visual Screening (NVS) for hallucination detection, and Adaptive Token-level Contrastive Decoding (ATCD) for hallucination mitigation. CATCH addresses issues related to visual defects that cause diminished fine-grained feature perception and cumulative hallucinations in open-ended scenarios. It is applicable to various visual question-answering tasks without requiring any specific data or prior knowledge, and generalizes robustly to new tasks without additional training, opening new possibilities for advancing LVLM in various challenging applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。