融合合成与人工偏好数据,提升视频细粒度描述质量
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
- 三阶段训练:筛选高质合成文本,融合人类偏好
- 在VDCSCORE上达到新SOTA,人类评估显著优于现有方法
- 推出轻量版Cockatiel-8B,兼顾性能与实用
视频细粒度描述(VDC)是视觉-语言对齐的关键任务,能生成复杂视频内容的精细描述。本文首次全面评估当前最先进方法,系统发现两大缺陷:对特定描述维度存在偏差,且与人类偏好不一致。为此提出Cockatiel,一种三阶段训练流程,融合合成数据与人工偏好以提升VDC性能。第一阶段,基于精心标注数据集构建评分器,筛选在细粒度视频-文本对齐和人类偏好上表现优异的合成描述;第二阶段,使用该筛选数据集训练Cockatiel-13B,融合多模型优势与人类偏好;第三阶段,从Cockatiel-13B进一步蒸馏出更易部署的Cockatiel-8B。大量定量与定性实验表明,本方法在维均衡的VDCSCORE上达到新最优,且在人类评估中显著超越主流模型。
原文摘要 · Abstract (English)
Video Detailed Captioning (VDC) is a crucial task for vision-language bridging, enabling fine-grained descriptions of complex video content. In this paper, we first comprehensively benchmark current state-of-the-art approaches and systematically identified two critical limitations: biased capability towards specific captioning aspect and misalignment with human preferences. To address these deficiencies, we propose Cockatiel, a novel three-stage training pipeline that ensembles synthetic and human-aligned training for improving VDC performance. In the first stage, we derive a scorer from a meticulously annotated dataset to select synthetic captions high-performing on certain fine-grained video-caption alignment and human-preferred while disregarding others. Then, we train Cockatiel-13B, using this curated dataset to infuse it with assembled model strengths and human preferences. Finally, we further distill Cockatiel-8B from Cockatiel-13B for the ease of usage. Extensive quantitative and qualitative experiments reflect the effectiveness of our method, as we not only set new state-of-the-art performance on VDCSCORE in a dimension-balanced way but also surpass leading alternatives on human preference by a large margin as depicted by the human evaluation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。