用文本奖励模型融合提升视觉语言模型的内容评估能力。
Transferring Textual Preferences to Vision-Language Understanding through Model Merging
- 通过模型融合将文本奖励模型与视觉语言模型结合
- 融合后在内容评估上优于原始模型和独立文本模型
- 无需训练,适合快速部署偏好评估功能
大型视觉语言模型(LVLMs)在多种多模态任务中表现优异。然而,其生成内容的评估能力仍有限,且使用偏好数据训练视觉语言奖励模型(VLRMs)成本高昂。本文探索了一种无需训练的替代方法:将基于文本的奖励模型(RMs)与LVLMs融合,以构建VLRMs。实验表明,该方法能有效提升模型对生成内容的评分性能,优于原始LVLMs及独立文本奖励模型,为将文本偏好高效注入视觉语言模型提供了新路径。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) perform outstandingly across various multimodal tasks. However, their ability to evaluate generated content remains limited, and training vision-language reward models (VLRMs) with preference data is computationally expensive. This paper explores a training-free alternative by merging text-based reward models (RMs) with LVLMs to create VLRMs. Our approach shows that integrating these models leads to improved performance over LVLMs' scoring and text-based RMs, offering an efficient method for incorporating textual preferences into LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。