用AI生成反馈数据,提升视觉语言模型的对齐效果。
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment

- 用现成模型生成8.2万条多模态指令和解释,无需人工标注。
- 模型在感知与认知任务上分别提升6.9%和9.5%,幻觉减少。
- 适合研究大模型对齐、AI反馈训练的开发者与研究人员。
随着大规模视觉语言模型(LVLM)快速发展,高质量、多样化的对齐数据需求日益迫切。然而,依赖人工标注的数据构建成本高、耗时长。本文探索利用AI反馈扩展监督信号以对齐LVLM的有效性。我们提出VLFeedback,首个大规模多模态反馈数据集,包含超过82,000条由现成模型生成的多模态指令及完整推理过程,无需人工标注。为评估其效果,我们基于该数据集,通过直接偏好优化微调得到Silkie模型。结果表明,该模型在帮助性、视觉忠实度与安全性方面表现优异,在感知与认知任务上分别优于基线模型6.9%和9.5%,在MMHal-Bench上显著降低幻觉问题,并具备更强抗红队攻击能力。分析还显示,AI反馈尤其有助于提升偏好多样性,带来更全面的性能改进。数据集、训练代码与模型已开源。
原文摘要 · Abstract (English)
As large vision-language models (LVLMs) evolve rapidly, the demand for high-quality and diverse data to align these models becomes increasingly crucial. However, the creation of such data with human supervision proves costly and time-intensive. In this paper, we investigate the efficacy of AI feedback to scale supervision for aligning LVLMs. We introduce VLFeedback, the first large-scale vision-language feedback dataset, comprising over 82K multi-modal instructions and comprehensive rationales generated by off-the-shelf models without human annotations. To evaluate the effectiveness of AI feedback for vision-language alignment, we train Silkie, an LVLM fine-tuned via direct preference optimization on VLFeedback. Silkie showcases exceptional performance regarding helpfulness, visual faithfulness, and safety metrics. It outperforms its base model by 6.9\% and 9.5\% in perception and cognition tasks, reduces hallucination issues on MMHal-Bench, and exhibits enhanced resilience against red-teaming attacks. Furthermore, our analysis underscores the advantage of AI feedback, particularly in fostering preference diversity to deliver more comprehensive improvements. Our dataset, training code and models are available at https://vlf-silkie.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。