arXiv:2512.04356cs.CVcs.AI2025-12被引 3

通过自增强对比对齐,减少多模态大模型在视频描述中的物体和动作幻觉。

Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment

  • 设计自增强方案识别并生成幻觉对抗样本
  • 引入轨迹-短语对比对齐,提升物体与动作匹配精度
  • 适用于需要高真实性的视频理解任务

多模态大模型在视频描述生成方面表现出色,但常出现事实性错误导致严重幻觉。现有工作主要针对静态图像的幻觉缓解,而动态视频中物体与动作的联合幻觉仍难解决。为此,我们提出自增强对比对齐(SANTA)框架,通过消除虚假关联、强化视觉事实来提升描述忠实度。SANTA采用幻觉自增强机制识别模型潜在幻觉,并将原描述转换为对比负例;同时设计轨迹-短语对比对齐策略,将区域物体与关系引导的动作与对应视觉和时间短语精准匹配。大量实验表明,SANTA在幻觉检测基准上优于现有方法,显著降低物体与动作幻觉。

原文摘要 · Abstract (English)

Recent advancement in multimodal LLMs (MLLMs) has demonstrated their remarkable capability to generate descriptive captions for input videos. However, these models suffer from factual inaccuracies in the generated descriptions, causing severe hallucination issues. While prior works have explored alleviating hallucinations for static images, jointly mitigating visual object and temporal action hallucinations for dynamic videos remains a challenging and unsolved task. To tackle this challenge, we propose a Self-Augmented Contrastive Alignment (SANTA) framework for enabling object and action faithfulness by exempting the spurious correlations and enforcing the emphasis on visual facts. SANTA employs a hallucinative self-augmentation scheme to identify the potential hallucinations that lie in the MLLM and transform the original captions to the contrasted negatives. Furthermore, we develop a tracklet-phrase contrastive alignment to match the regional objects and relation-guided actions with their corresponding visual and temporal phrases. Extensive experiments demonstrate that SANTA outperforms existing methods in alleviating object and action hallucinations, yielding superior performance on the hallucination examination benchmarks.

多模态大模型幻觉抑制视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。