arXiv:2411.16201cs.LGcs.CL2024-11

用AI反馈自动构建视频问答偏好数据,提升多模态大模型对齐能力

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models

  • 通过多AI生成响应并评分,自动构建高质量视频问答偏好数据集
  • 提出迭代式弱到强强化学习框架,充分挖掘数据中的对齐信息
  • 适合研究多模态大模型对齐、自动数据构建的学者与开发者

高质量视频文本偏好数据对多模态大语言模型(MLLMs)对齐至关重要,但现有数据稀缺。人工标注成本高且不可靠,而温度控制的AI生成响应缺乏多样性。为此,本文提出一个名为MMAIP-V的高质量视频问答偏好数据集,通过从响应分布中采样并使用外部评分函数评估生成质量。为进一步利用该数据,提出迭代式弱到强强化学习框架Iter-W2S-RLAIF,通过不断更新参考模型并进行参数外推,逐步增强MLLMs的对齐能力。此外,设计了无偏且信息完整的视频问答评估方案。实验表明,MMAIP-V有助于MLLMs的偏好学习,Iter-W2S-RLAIF能充分挖掘其对齐潜力。该基于AI反馈的自动化数据生成流程可显著推动未来多模态大模型对齐研究。代码与数据集已公开。

原文摘要 · Abstract (English)

High-quality video-text preference data is crucial for Multimodal Large Language Models (MLLMs) alignment. However, existing preference data is very scarce. Obtaining VQA preference data for preference training is costly, and manually annotating responses is highly unreliable, which could result in low-quality pairs. Meanwhile, AI-generated responses controlled by temperature adjustment lack diversity. To address these issues, we propose a high-quality VQA preference dataset, called \textit{\textbf{M}ultiple \textbf{M}ultimodal \textbf{A}rtificial \textbf{I}ntelligence \textbf{P}reference Datasets in \textbf{V}QA} (\textbf{MMAIP-V}), which is constructed by sampling from the response distribution set and using an external scoring function for response evaluation. Furthermore, to fully leverage the preference knowledge in MMAIP-V and ensure sufficient optimization, we propose \textit{\textbf{Iter}ative \textbf{W}eak-to-\textbf{S}trong \textbf{R}einforcement \textbf{L}earning from \textbf{AI} \textbf{F}eedback for video MLLMs} (\textbf{Iter-W2S-RLAIF}), a framework that gradually enhances MLLMs' alignment capabilities by iteratively updating the reference model and performing parameter extrapolation. Finally, we propose an unbiased and information-complete evaluation scheme in VQA evaluation. Experiments demonstrate that MMAIP-V is beneficial for MLLMs in preference learning and Iter-W2S-RLAIF fully exploits the alignment information in MMAIP-V. We believe that the proposed automatic VQA preference data generation pipeline based on AI feedback can greatly promote future work in the MLLMs alignment. \textbf{Code and dataset are available} \href{https://anonymous.4open.science/r/MMAIP-V_Iter-W2S-RLAIF-702F}{MMAIP-V\_Iter-W2S-RLAIF-702F}.

多模态偏好学习AI反馈数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。