arXiv:2508.12458cs.CL2025-08被引 1

用模型自生成数据优化视觉指令跟随,提升效率与效果。

M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following

  • 基于多模态对齐与自一致性评分,自动筛选高质量训练样本。
  • 在多个基准上超越传统方法,7B/13B模型均表现更优。
  • 适合追求高效微调的视觉语言模型研究者与开发者。

大型视觉语言模型(LVLMs)在复杂多模态指令遵循任务中潜力巨大,但其训练常受限于人工标注成本高、不一致的问题。传统监督微调(SFT)及现有偏好优化方法如RLHF和DPO难以高效利用模型自身生成空间来发现高信息量的“困难负例”样本。为此,我们提出多模态模型引导的偏好优化(M3PO),一种数据高效的新型方法,旨在提升LVLM在视觉指令遵循上的能力。M3PO从多样化的LVLM生成候选样本池中智能选择最具“学习价值”的偏好样本对,该选择基于双重信号:外部质量评估的多模态对齐分数(MAS)和内部信念度量的自一致性/置信度(对数概率)。两者融合形成新颖的M3P得分,精准识别出模型可能自信生成却错误的优选与难判负选响应。这些高质量偏好对随后用于对基础LVLM(如LLaVA-1.5 7B/13B)进行高效直接偏好优化(DPO)微调,采用LoRA技术。大量实验表明,M3PO在涵盖多模态指令遵循的综合基准(MME-Bench、POPE、IFT、Human Pref. Score)上持续优于强基线,包括SFT、模拟RLHF、原始DPO和RM-DPO。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) hold immense potential for complex multimodal instruction following, yet their development is often hindered by the high cost and inconsistency of human annotation required for effective fine-tuning and preference alignment. Traditional supervised fine-tuning (SFT) and existing preference optimization methods like RLHF and DPO frequently struggle to efficiently leverage the model's own generation space to identify highly informative "hard negative" samples. To address these challenges, we propose Multimodal-Model-Guided Preference Optimization (M3PO), a novel and data-efficient method designed to enhance LVLMs' capabilities in visual instruction following. M3PO intelligently selects the most "learning-valuable" preference sample pairs from a diverse pool of LVLM-generated candidates. This selection is driven by a sophisticated mechanism that integrates two crucial signals: a Multimodal Alignment Score (MAS) to assess external quality and the model's Self-Consistency / Confidence (log-probability) to gauge internal belief. These are combined into a novel M3P-Score, which specifically identifies preferred responses and challenging dispreferred responses that the model might confidently generate despite being incorrect. These high-quality preference pairs are then used for efficient Direct Preference Optimization (DPO) fine-tuning on base LVLMs like LLaVA-1.5 (7B/13B) using LoRA. Our extensive experiments demonstrate that M3PO consistently outperforms strong baselines, including SFT, simulated RLHF, vanilla DPO, and RM-DPO, across a comprehensive suite of multimodal instruction following benchmarks (MME-Bench, POPE, IFT, Human Pref. Score).

视觉语言模型偏好优化数据效率多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。