用排序优化提升视频文本对齐质量,更精准评估生成内容与视觉的契合度。
Learning to Rank Caption Chains for Video-Text Alignment
- 通过重复降级生成有序的长视频描述链,实现大规模排序训练。
- 排序优化在长文本生成与评估中优于传统二元偏好方法。
- 需微调视觉编码器,说明该方法不只是语言模型重加权。
直接偏好优化(DPO)虽能有效训练语言模型生成优选响应,但其二元‘胜者全得’机制对视觉语言模型不适用,因响应质量高度依赖视觉内容。即使某响应不如另一选项,仍可能忠实于视觉输入。标准布拉德利-特里尔DPO忽略此细微差别,过度强调胜出响应而忽视失败响应的视觉保真度。本文提出以排序优化替代,更精确衡量响应与视觉内容的契合度。聚焦使用详细视频描述进行视频-文本对齐,通过反复降级生成大规模、完全有序的描述链。实验表明,排序优化在长文本生成与评估中优于二元DPO;更重要的是,我们发现此类方法需微调视觉编码器才有效,挑战了DPO仅为语言重加权的观点。
原文摘要 · Abstract (English)
Direct preference optimization (DPO) is an effective technique to train language models to generate preferred over dispreferred responses. However, this binary "winner-takes-all" approach is suboptimal for vision-language models whose response quality is highly dependent on visual content. In particular, a response may still be faithful to the visual inputs even if it is less preferable than an alternative. The standard Bradley-Terry DPO formulation lacks this nuance, upweighting winning responses without sufficient regard for whether the "losing" response still maintains high visual fidelity. In this work, we investigate ranking optimization as an alternative that more precisely situates responses' faithfulness to visual inputs. We focus on video-text alignment using detailed video captions, proposing a method to generate challenging, totally ordered caption chains at scale through repeated caption degradation. Our results show ranking optimization outperforms binary DPO for long-form content generation and assessment, and importantly, we find that these approaches require finetuning of the vision encoder to be effective, challenging the view of DPO as purely a language-reweighting process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。