平衡视频描述中的动作与细节,提升生成质量
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
- 构建新数据集HMD-270K,融合动作与细节信息
- 引入新奖励机制CSER,提升描述完整性和准确性
- 适合关注视频理解与多模态生成的研究者
视频字幕生成旨在生成全面且连贯的视频内容描述,推动视频理解和生成的发展。然而,现有方法常因动作与细节失衡而表现不佳,模型过度侧重某一方面而忽略另一方,导致描述不完整,影响视频理解与生成的一致性。为此,我们从数据和优化两方面提出解决方案:1)数据层面,通过两阶段管道(运动-细节融合与细粒度检验)构建了包含27万样本的和谐动作-细节数据集HMD-270K;2)优化层面,提出基于组相对策略优化(GRPO)的字幕集合等价奖励(CSER),通过单元到集合匹配与双向验证,增强对动作与细节的捕捉。基于HMD-270K的监督微调与结合CSER的GRPO后训练,我们开发出OwlCap——一个具备动作-细节平衡能力的多模态大语言模型。实验表明,相比基线模型,OwlCap在两个基准上均有显著提升:在注重细节的VDC数据集上准确率提高4.2,在注重动作的DREAM-1K数据集上F1值提升4.6。HMD-270K数据集与OwlCap模型将公开发布,以促进该领域研究进展。
原文摘要 · Abstract (English)
Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance, as models tend to overemphasize one aspect while neglecting the other. This imbalance results in incomplete captions, which in turn leads to a lack of consistency in video understanding and generation. To address this issue, we propose solutions from two aspects: 1) Data aspect: We constructed the Harmonizing Motion-Detail 270K (HMD-270K) dataset through a two-stage pipeline: Motion-Detail Fusion (MDF) and Fine-Grained Examination (FGE). 2) Optimization aspect: We introduce the Caption Set Equivalence Reward (CSER) based on Group Relative Policy Optimization (GRPO). CSER enhances completeness and accuracy in capturing both motion and details through unit-to-set matching and bidirectional validation. Based on the HMD-270K supervised fine-tuning and GRPO post-training with CSER, we developed OwlCap, a powerful video captioning multi-modal large language model (MLLM) with motion-detail balance. Experimental results demonstrate that OwlCap achieves significant improvements compared to baseline models on two benchmarks: the detail-focused VDC (+4.2 Acc) and the motion-focused DREAM-1K (+4.6 F1). The HMD-270K dataset and OwlCap model will be publicly released to facilitate video captioning research community advancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。