arXiv:2509.24491cs.CVcs.AI2025-09被引 1

通过语义课程优化,显著减少多模态大模型的视觉幻觉问题。

Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

  • 构建由易到难的语义对比数据集,分阶段训练模型理解图文一致性。
  • 在多个基准上将幻觉率降低62.9%,同时保持模型通用能力稳定。
  • 首次融合语义、对称性与课程学习,适合需高准确性的视觉问答场景。

多模态大语言模型(MLLMs)在各类任务中表现显著提升,但仍面临视觉幻觉这一关键问题——生成内容与视觉证据矛盾。尽管直接偏好优化(DPO)广泛用于对齐,但其在MLLM中常难以捕捉细粒度语义差异,且易导致捷径学习。为此,我们提出语义课程偏好优化(SCPO),一种新型的MLLM对齐框架。SCPO基于自建的语义课程偏好对数据集,按难度排序提供细粒度语义对比,采用动态参考模型与新颖的对称双向目标,实现文本与视觉偏好同步学习。据我们所知,SCPO是首个统一语义、对称性与课程学习的MLLM对齐框架,能有效缓解视觉幻觉。在多种规模和版本的LLaVA模型上进行的大量实验表明,相比基线模型,SCPO在多个幻觉基准上表现更优,幻觉率最高降低62.9%。此外,在泛化基准上的评估显示,SCPO在提升事实准确性的同时保持通用能力,性能在多个通用视觉-语言基准上保持稳定。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have significantly improved the performance of various tasks, but continue to suffer from visual hallucinations, a critical issue where generated responses contradict visual evidence. While Direct Preference Optimization(DPO) is widely used for alignment, its application to MLLMs often fails to capture fine-grained semantic differences and encourages shortcut learning. To address these challenges, we propose Semantic Curriculum Preference Optimization (SCPO), a novel framework for MLLM alignment. SCPO employs a progressive, easy-to-hard curriculum built upon our Semantic Curriculum Preference Pairs dataset, which provides fine-grained semantic contrasts sorted by difficulty. This curriculum is trained with a dynamic reference model and a novel symmetric, bidirectional objective to facilitate simultaneous learning from both textual and visual preferences. To our knowledge, SCPO is the first framework to unify semantics, symmetry, and curriculum for MLLMs alignment, effectively mitigating visual hallucinations. Extensive experiments on LLaVA models across various scales and versions validate that SCPO demonstrates superior performance compared to baseline models on multiple hallucination benchmarks, reducing the hallucination rate by up to 62.9%. Moreover, evaluations on generalized benchmarks show that SCPO improves factuality while preserving general capabilities, with its performance remaining stable across general vision-language benchmarks.

多模态模型幻觉抑制模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。