通过模型引导的反事实数据,精准减少视频大模型幻觉。
MACD: Model-Aware Contrastive Decoding via Counterfactual Data
- 用模型自身反馈定位易引发幻觉的物体区域,生成针对性反事实输入。
- 在多个数据集上显著降低幻觉率,小物体、遮挡物场景提升明显。
- 适用于Qwen、InternVL等多类视频大模型,不牺牲准确率。
视频语言模型(Video-LLMs)在视觉证据弱、模糊或有偏时容易产生幻觉,生成看似合理但无依据的内容。现有方法如对比解码(CD)依赖随机扰动生成对比数据,但常无法精准针对诱发幻觉的视觉线索或模型弱点。本文提出模型感知的反事实数据对比解码(MACD),结合模型引导的反事实构造与对比解码。MACD利用视频大模型自身反馈,识别导致幻觉的关键物体区域,生成针对物体级别的反事实输入,而非随意的帧或时间修改。这些反事实输入被融入对比解码,以强制解码过程中选择有证据支持的词元。在EventHallusion、MVBench、Perception-test和Video-MME等多个数据集上的实验表明,MACD能持续降低幻觉率,同时保持或提升任务准确率,对小物体、遮挡物或共现物体场景尤其有效,适用于Qwen、InternVL等多种视频大模型。
原文摘要 · Abstract (English)
Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous, or biased. Existing methods, such as contrastive decoding (CD), rely on random perturbations to construct contrastive data for hallucination mitigation, but often fail to target the visual cues that drive hallucination or align with model weaknesses. We propose Model-Aware Counterfactual Data based Contrastive Decoding (MACD), an inference strategy that combines model-guided counterfactual construction with contrastive decoding. MACD uses the Video-LLM's own feedback to identify object regions most responsible for hallucination, generating targeted object-level counterfactual inputs rather than arbitrary frame or temporal modifications. These counterfactual inputs are integrated into CD to enforce evidence-grounded token selection during decoding. Experiments on EventHallusion, MVBench, Perception-test, and Video-MME show that MACD consistently reduces hallucination while maintaining or improving task accuracy across diverse Video-LLMs, including Qwen and InternVL, with especially strong gains in scenarios involving small, occluded, or co-occurring objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。