解决多图理解中模型幻觉问题,通过分层优化提升视觉细节感知
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
- 分两层优化:全局上下文与局部细节,减少认知偏差
- 在多图任务上显著降低幻觉率,提升理解准确性
- 适合需要精准多图分析的场景,如医疗影像、复杂图表解析
多模态大语言模型在单图任务中表现优异,但在多图理解时因跨模态对齐问题易产生幻觉(遗漏上下文、混淆信息、误解读)。现有基于直接偏好优化(DPO)的方法仅以单张图像为参考,忽略整体上下文建模。本文提出上下文到线索的直接偏好优化(CcDPO),一种多层级偏好优化框架,通过从序列上下文聚焦到局部视觉线索,增强多图场景下的图像感知能力。其包含:(i) 上下文层优化:重新评估多图上下文理解中的认知偏差,并引入低成本全局序列偏好以缓解偏差;(ii) 针线层优化:通过区域定向视觉提示和多模态偏好监督,引导关注细粒度视觉细节。为支持可扩展优化,我们构建了多尺度-42k(MultiScope-42k)数据集,包含高质量多层级偏好对。实验表明,CcDPO显著减少幻觉现象,并在通用单图与多图任务中实现一致性能提升。
原文摘要 · Abstract (English)
Multi-modal Large Language Models (MLLMs) excel at single-image tasks but struggle with multi-image understanding due to cross-modal misalignment, leading to hallucinations (context omission, conflation, and misinterpretation). Existing methods using Direct Preference Optimization (DPO) constrain optimization to a solitary image reference within the input sequence, neglecting holistic context modeling. We propose Context-to-Cue Direct Preference Optimization (CcDPO), a multi-level preference optimization framework that enhances per-image perception in multi-image settings by zooming into visual clues -- from sequential context to local details. It features: (i) Context-Level Optimization : Re-evaluates cognitive biases underlying MLLMs' multi-image context comprehension and integrates a spectrum of low-cost global sequence preferences for bias mitigation. (ii) Needle-Level Optimization : Directs attention to fine-grained visual details through region-targeted visual prompts and multimodal preference supervision. To support scalable optimization, we also construct MultiScope-42k, an automatically generated dataset with high-quality multi-level preference pairs. Experiments show that CcDPO significantly reduces hallucinations and yields consistent performance gains across general single- and multi-image tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。