通过对称偏好学习,更精准地减少多模态大模型的幻觉。
Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- 设计对称偏好优化框架,直接利用回答对进行监督。
- 在五个基准上显著降低幻觉率,提升视觉理解能力。
- 适合关注多模态模型可靠性与生成准确性的研究者。
直接偏好优化(DPO)已成为缓解多模态大语言模型(MLLMs)幻觉的有效方法。尽管现有方法通过视觉导向的对比目标增强了模型对视觉输入的关注,从而减少了幻觉,但仍存在优化目标不严谨和偏好监督间接的问题。为此,我们提出对称多模态偏好优化(SymMPO),采用对称偏好学习并结合直接偏好监督(即回答对),以增强视觉理解,同时保持与标准DPO严格的理论一致性。除传统序数偏好学习外,SymMPO引入偏好间隙一致性损失,定量调节对称偏好对之间的偏好差距。在五个基准上的综合评估验证了SymMPO在缓解MLLM幻觉方面的优越性能。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visual inputs and hence reducing hallucination, they suffer from non-rigorous optimization objective function and indirect preference supervision. To address these limitations, we propose a Symmetric Multimodal Preference Optimization (SymMPO), which conducts symmetric preference learning with direct preference supervision (i.e., response pairs) for visual understanding enhancement, while maintaining rigorous theoretical alignment with standard DPO. In addition to conventional ordinal preference learning, SymMPO introduces a preference margin consistency loss to quantitatively regulate the preference gap between symmetric preference pairs. Comprehensive evaluation across five benchmarks demonstrate SymMPO's superior performance, validating its effectiveness in hallucination mitigation of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。