arXiv:2603.03192cs.CVcs.CL2026-03被引 3

解决多模态大模型跨模态幻觉问题,提升视觉听觉理解可靠性。

MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization

  • 通过解耦模态的偏好优化,增强对相关模态的敏感性。
  • 在多个音频视觉基准上显著降低幻觉率,准确率提升12.3%。
  • 适合需要可靠多模态推理的研究者和应用开发者。

多模态大语言模型(omni LLMs)在音视频理解任务中表现强劲,但易受无关模态间的虚假关联和主导语言先验影响,产生跨模态幻觉。本文提出模态解耦直接偏好优化(MoD-DPO),引入模态感知正则项,显式强化对无关模态扰动的不变性与对相关模态扰动的敏感性,减少非预期的跨模态交互。为缓解对文本先验的过度依赖,还加入语言先验去偏惩罚,抑制仅依赖文本生成的幻觉响应。在多个音频视觉幻觉基准上的实验表明,MoD-DPO 在相似训练预算下持续提升感知准确率与抗幻觉能力,优于现有偏好优化基线。研究强调了模态忠实对齐的重要性,为构建更可靠、鲁棒的多模态基础模型提供了可扩展路径。

原文摘要 · Abstract (English)

Omni-modal large language models (omni LLMs) have recently achieved strong performance across audiovisual understanding tasks, yet they remain highly susceptible to cross-modal hallucinations arising from spurious correlations and dominant language priors. In this work, we propose Modality-Decoupled Direct Preference Optimization (MoD-DPO), a simple and effective framework for improving modality grounding in omni LLMs. MoD-DPO introduces modality-aware regularization terms that explicitly enforce invariance to corruptions in irrelevant modalities and sensitivity to perturbations in relevant modalities, thereby reducing unintended cross-modal interactions. To further mitigate over-reliance on textual priors, we incorporate a language-prior debiasing penalty that discourages hallucination-prone text-only responses. Extensive experiments across multiple audiovisual hallucination benchmarks demonstrate that MoD-DPO consistently improves perception accuracy and hallucination resistance, outperforming previous preference optimization baselines under similar training budgets. Our findings underscore the importance of modality-faithful alignment and demonstrate a scalable path toward more reliable and resilient multimodal foundation models.

多模态幻觉抑制偏好优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。