让多模态模型更懂指令,提升理解能力
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
- 基于指令完成度自动构建偏好数据,避免人为偏差
- 在多个评测集上显著提升视觉问答与文本理解表现
- 适合需要精准理解指令的多模态应用开发者
偏好对齐已成为提升多模态大模型(MLLMs)性能的有效策略。现有方法主要关注幻觉抑制,却忽视了多模态理解能力的关键因素,导致改进局限于幻觉缓解。为此,我们提出指令导向偏好对齐(IPA),一种可扩展的框架,能基于指令完成度自动构建对齐偏好。该方法结合自动化偏好生成与专用验证流程,识别指令相关因素,减少响应表征的不稳定性。此外,IPA引入渐进式偏好收集管道,通过模型自演化和参考引导优化,召回更具挑战性的样本。在Qwen2VL-7B上的实验表明,IPA在多个基准测试中均表现优异,涵盖幻觉评估、视觉问答及文本理解任务,充分展现了其提升通用理解能力的潜力。
原文摘要 · Abstract (English)
Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly target hallucination factors, they overlook the factors essential for multi-modal comprehension capabilities, often narrowing their improvements on hallucination mitigation. To bridge this gap, we propose Instruction-oriented Preference Alignment (IPA), a scalable framework designed to automatically construct alignment preferences grounded in instruction fulfillment efficacy. Our method involves an automated preference construction coupled with a dedicated verification process that identifies instruction-oriented factors, avoiding significant variability in response representations. Additionally, IPA incorporates a progressive preference collection pipeline, further recalling challenging samples through model self-evolution and reference-guided refinement. Experiments conducted on Qwen2VL-7B demonstrate IPA's effectiveness across multiple benchmarks, including hallucination evaluation, visual question answering, and text understanding tasks, highlighting its capability to enhance general comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。