用视觉编码器做细粒度反馈,提升多模态模型对齐效果
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
- 利用视觉编码器生成逐令牌反馈,实现模型自对齐
- 在多个数据集上超越依赖外部数据的偏好训练方法
- 适合研究多模态对齐与安全生成的学者使用
大型语言模型(LLMs)和预训练视觉模型的发展推动了视觉-语言大模型(VLLMs)的进步,增强了视觉与语言模态间的交互。尽管在多个领域取得显著成果,VLLMs仍面临模态对齐难题,易导致幻觉和不安全内容生成。现有对齐方法通常依赖粗粒度反馈和外部数据集,限制了可扩展性和性能。本文提出FiSAO(细粒度自对齐优化),一种新颖的自对齐方法,利用模型自身的视觉编码器作为细粒度验证器,在无需额外数据的情况下提升视觉-语言对齐。通过视觉编码器提供的逐令牌反馈,FiSAO显著改善对齐效果,甚至优于需额外数据的传统偏好调优方法。理论分析与实验验证表明,FiSAO有效解决了VLLMs中的对齐偏差问题,是首个将逐令牌奖励应用于此类模型的工作。
原文摘要 · Abstract (English)
The recent advancements in large language models (LLMs) and pre-trained vision models have accelerated the development of vision-language large models (VLLMs), enhancing the interaction between visual and linguistic modalities. Despite their notable success across various domains, VLLMs face challenges in modality alignment, which can lead to issues like hallucinations and unsafe content generation. Current alignment techniques often rely on coarse feedback and external datasets, limiting scalability and performance. In this paper, we propose FiSAO (Fine-Grained Self-Alignment Optimization), a novel self-alignment method that utilizes the model's own visual encoder as a fine-grained verifier to improve vision-language alignment without the need for additional data. By leveraging token-level feedback from the vision encoder, FiSAO significantly improves vision-language alignment, even surpassing traditional preference tuning methods that require additional data. Through both theoretical analysis and experimental validation, we demonstrate that FiSAO effectively addresses the misalignment problem in VLLMs, marking the first instance of token-level rewards being applied to such models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。