用新方法提升视频细粒度描述生成质量,更准更快。
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- 结合视觉语言模型与大模型构建偏好数据对,兼顾成本与质量。
- 在多个视频描述数据集上显著优于DPO,训练效率提升20%。
- 适合追求高质量视频生成与高效训练的科研与工程人员。
细粒度视频字幕生成旨在生成详细且时间连贯的视频内容描述,但现有方法难以捕捉细微动态和丰富细节。本文利用偏好学习增强视觉-语言模型在细粒度视频字幕任务中的表现,同时缓解直接偏好优化(DPO)的若干局限。首先,提出一种基于视觉语言模型内在特性并辅以大语言模型部分协助的偏好对构建流程,实现成本与数据质量的最佳平衡。其次,提出协同偏好优化(SynPO),该方法相比DPO及其变体具有显著优势:防止负向偏好主导优化过程,显式保留模型语言能力以避免目标偏离,并通过无需参考模型提升训练效率。我们在视频字幕基准(如VDC、VDD、VATEX)及通用NLP任务(包括语言理解与偏好评估)上广泛评估SynPO,使用多种预训练模型。结果表明,SynPO持续优于DPO变体,且训练效率提升20%。代码已开源。
原文摘要 · Abstract (English)
Fine-grained video captioning aims to generate detailed, temporally coherent descriptions of video content. However, existing methods struggle to capture subtle video dynamics and rich detailed information. In this paper, we leverage preference learning to enhance the performance of vision-language models in fine-grained video captioning, while mitigating several limitations inherent to direct preference optimization (DPO). First, we propose a pipeline for constructing preference pairs that leverages the intrinsic properties of VLMs along with partial assistance from large language models, achieving an optimal balance between cost and data quality. Second, we propose Synergistic Preference Optimization (SynPO), a novel optimization method offering significant advantages over DPO and its variants. SynPO prevents negative preferences from dominating the optimization, explicitly preserves the model's language capability to avoid deviation of the optimization objective, and improves training efficiency by eliminating the need for the reference model. We extensively evaluate SynPO not only on video captioning benchmarks (e.g., VDC, VDD, VATEX) but also across well-established NLP tasks, including general language understanding and preference evaluation, using diverse pretrained models. Results demonstrate that SynPO consistently outperforms DPO variants while achieving 20\% improvement in training efficiency. Code is available at https://github.com/longmalongma/SynPO
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。