构建20万条数据集,让多模态大模型更懂人类偏好。
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
- 构建20万条多样化图文数据,覆盖复杂问题与多格式回答。
- 微调后模型在人类偏好评估中显著提升,同时保持VQA性能。
- 适合研究多模态对齐、人机价值观一致性的学者使用。
开源多模态大模型(MLLM)近年发展聚焦于基础能力提升,但在人类偏好对齐方面仍存在显著空白。本文提出OmniAlign-V,一个包含20万条高质量训练样本的综合性数据集,涵盖多样图像、复杂问题及多种响应格式,旨在提升MLLM与人类偏好的对齐程度。同时,我们构建了MM-AlignBench,一个由人工标注的基准测试集,专门用于评估模型在人类价值观对齐方面的表现。实验表明,采用监督微调(SFT)或直接偏好优化(DPO)方法,用OmniAlign-V对MLLM进行微调,可显著增强其与人类偏好的一致性,同时在标准VQA基准上保持或提升性能,有效保留模型的基础能力。相关数据集、基准、代码与模型权重已公开于https://github.com/PhoenixZ810/OmniAlign-V。
原文摘要 · Abstract (English)
Recent advancements in open-source multi-modal large language models (MLLMs) have primarily focused on enhancing foundational capabilities, leaving a significant gap in human preference alignment. This paper introduces OmniAlign-V, a comprehensive dataset of 200K high-quality training samples featuring diverse images, complex questions, and varied response formats to improve MLLMs' alignment with human preferences. We also present MM-AlignBench, a human-annotated benchmark specifically designed to evaluate MLLMs' alignment with human values. Experimental results show that finetuning MLLMs with OmniAlign-V, using Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO), significantly enhances human preference alignment while maintaining or enhancing performance on standard VQA benchmarks, preserving their fundamental capabilities. Our datasets, benchmark, code and checkpoints have been released at https://github.com/PhoenixZ810/OmniAlign-V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。