arXiv:2411.14797cs.LGcs.AI2024-11CVPR被引 6

用负样本监督替代多模态强化学习,高效对齐视觉语言模型

Continual SFT Matches Multimodal RLHF with Negative Supervision

  • 提出nSFT方法,从拒绝响应中提取负向监督信号
  • 在多个数据集和模型上达到与RLHF相当的对齐效果
  • 仅需单个模型,内存占用远低于传统RLHF方法

多模态强化学习人类反馈(RLHF)通常在监督微调(SFT)后进行,以持续提升视觉语言模型(VLM)的理解能力。传统观点认为其在偏好对齐阶段优于持续SFT。本文发现,多模态RLHF的核心价值在于负向监督——即被拒绝响应的逻辑值。为此,我们提出一种新的负向监督微调(nSFT)方法,充分挖掘该信息。nSFT解耦了RLHF范式中的负向监督,并通过简单的SFT损失持续对齐模型。相比传统多模态RLHF(如DPO需2个、PPO需4个大型VLM),nSFT更节省内存。我们在不同数据源、基线VLM和评估指标下,严格对比了nSFT与多种多模态RLHF方法,验证了其有效性。丰富的消融实验也支持了我们的假设。希望本工作能推动大视觉语言模型的合理对齐研究。

原文摘要 · Abstract (English)

Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this preference alignment stage. In this paper, we observe that the inherent value of multimodal RLHF lies in its negative supervision, the logit of the rejected responses. We thus propose a novel negative supervised finetuning (nSFT) approach that fully excavates these information resided. Our nSFT disentangles this negative supervision in RLHF paradigm, and continually aligns VLMs with a simple SFT loss. This is more memory efficient than multimodal RLHF where 2 (e.g., DPO) or 4 (e.g., PPO) large VLMs are strictly required. The effectiveness of nSFT is rigorously proved by comparing it with various multimodal RLHF approaches, across different dataset sources, base VLMs and evaluation metrics. Besides, fruitful of ablations are provided to support our hypothesis. We hope this paper will stimulate further research to properly align large vision language models.

视觉语言模型负向监督SFT优化高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。