arXiv:2602.00559cs.CVcs.AI2026-02被引 1

提出新方法对抗视频多模态模型的复合幻觉问题。

Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models

  • 设计三路校准解码框架,动态干扰并强化视觉证据。
  • 在39个模型上测试,性能提升超10%。
  • 适合研究多模态大模型幻觉与推理的学者。

当前视频幻觉缓解研究主要聚焦单一错误类型,而由多个时空因素交互导致的复合幻觉仍缺乏系统探索。本文提出OmniVCHall基准,用于系统评估视频多模态大语言模型(VLLMs)中的孤立与复合幻觉。该基准涵盖多样视频领域,引入新型基于摄像头的幻觉类型,并建立细粒度分类体系,配备对抗性答案选项(如“全部正确”和“以上皆非”),防止捷径推理。对39个代表性VLLM的评估显示,即使先进模型(如Qwen3-VL和GPT-5)也出现显著性能下降。为此,我们提出TriCD对比解码框架,包含三路径校准机制:自适应扰动控制器动态选择干扰操作生成负样本视频,显著性引导增强模块自适应强化词级视觉证据。上述组件通过强化学习优化,在复合幻觉场景下促进精准决策。实验结果表明,TriCD在两个代表性骨干模型上均实现稳定提升,平均准确率提高超过10%。数据与代码见https://github.com/BMRETURN/OmniVCHall。

原文摘要 · Abstract (English)

Current research on video hallucination mitigation primarily focuses on isolated error types, leaving compositional hallucinations, arising from incorrect reasoning over multiple interacting spatial and temporal factors largely underexplored. We introduce OmniVCHall, a benchmark designed to systematically evaluate both isolated and compositional hallucinations in video multimodal large language models (VLLMs). OmniVCHall spans diverse video domains, introduces a novel camera-based hallucination type, and defines a fine-grained taxonomy, together with adversarial answer options (e.g., "All are correct" and "None of the above") to prevent shortcut reasoning. The evaluations of 39 representative VLLMs reveal that even advanced models (e.g., Qwen3-VL and GPT-5) exhibit substantial performance degradation. We propose TriCD, a contrastive decoding framework with a triple-pathway calibration mechanism. An adaptive perturbation controller dynamically selects distracting operations to construct negative video variants, while a saliency-guided enhancement module adaptively reinforces grounded token-wise visual evidences. These components are optimized via reinforcement learning to encourage precise decision-making under compositional hallucination settings. Experimental results show that TriCD consistently improves performance across two representative backbones, achieving an average accuracy improvement of over 10%. The data and code can be find at https://github.com/BMRETURN/OmniVCHall.

视频多模态幻觉抑制解码策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。