arXiv:2509.00723cs.AIcs.MM2025-09被引 7

解决多模态大模型幻觉问题,提升视觉音频理解能力

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination

论文配图:OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
图 1 · 摘自论文原文
  • 构建文本偏好样本对,增强音视频交互理解
  • 设计多模态偏好样本对,强化视听信息关注
  • 适用于需跨模态推理的多模态大模型优化

近年来,多模态大语言模型(OLLMs)在音视频理解与实时环境感知等任务中取得显著进展,但幻觉问题仍存在。与双模态类似,文本模态的先验倾向导致模型过度依赖文本线索而忽视视觉和音频信息。此外,全多模态场景引入新挑战:现有模型在训练中独立对齐视觉或听觉模态与文本,忽略视频与其对应音频的内在关联,导致在需要解析视频中隐藏音频线索时产生幻觉。为此,我们提出OmniDPO,一种用于缓解OLLM多模态幻觉的偏好对齐框架。具体包括:(1) 构建文本偏好样本对,以增强模型对音视频交互的理解;(2) 构建多模态偏好样本对,以加强模型对视觉与听觉信息的关注。实验表明,OmniDPO有效提升了多模态对齐能力,显著减少幻觉并增强跨模态推理性能。所有代码与数据集将在论文接收后公开。

原文摘要 · Abstract (English)

Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text modality tend to dominate, leading OLLMs to rely more heavily on textual cues while neglecting visual and audio information. In addition, fully multimodal scenarios introduce new challenges. Most existing models align visual or auditory modalities with text independently during training, while ignoring the intrinsic correlations between video and its corresponding audio. This oversight results in hallucinations when reasoning requires interpreting hidden audio cues embedded in video content. To address these challenges, we propose OmniDPO, a preference-alignment framework designed to mitigate hallucinations in OLLMs. Specifically, OmniDPO incorporates two strategies: (1) constructing text-preference sample pairs to enhance the model's understanding of audio-video interactions; and (2) constructing multimodal-preference sample pairs to strengthen the model's attention to visual and auditory information. By tackling both challenges, OmniDPO effectively improves multimodal grounding and reduces hallucination. Experiments conducted on two OLLMs demonstrate that OmniDPO not only effectively mitigates multimodal hallucinations but also significantly enhances the models' reasoning capabilities across modalities. All code and datasets will be released upon paper acceptance.

多模态幻觉抑制偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。