arXiv:2507.11336cs.CV2025-07被引 16

针对短视频的多模态细粒度描述,构建新数据集与轻量模型。

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

  • 设计三阶段人工闭环标注流程,平衡音视频模态。
  • 包含1000段TikTok视频与4000个问答对,支持跨模态理解评测。
  • 30亿参数模型通过两阶段训练,适配小样本高效生成。

真实用户生成视频(如TikTok)通常包含丰富的音视频交织内容。现有视频字幕基准与模型仍以视觉为主,忽视了音频在传达场景动态、说话人意图和叙事背景中的关键作用。为此,我们提出UGC-VideoCap数据集与模型框架,专为短时用户生成视频的细粒度多模态字幕设计。该数据集包含1000段TikTok视频,通过三阶段人工闭环标注流程,覆盖仅音频、仅视觉及音视频联合语义。同时提供4000个精心设计的问答对,用于测试单模态与跨模态理解能力。我们还提出30亿参数的UGC-VideoCaptioner(3B)模型,基于Gemini 2.5 Flash蒸馏而来。采用监督微调后接分组相对策略优化(GRPO)的两阶段训练策略,在有限数据下实现高效适配,保持竞争力。整体方案为非约束环境下多模态视频理解提供了高质量基础与数据高效解决方案。

原文摘要 · Abstract (English)

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric, overlooking the crucial role of audio in conveying scene dynamics, speaker intent, and narrative context. This lack of omni datasets and lightweight, capable models hampers progress in fine grained, multimodal video understanding. To address these challenges, we introduce UGC-VideoCap, a new benchmark and model framework specifically designed for detailed omnimodal captioning of short form user-generated videos. Unlike prior datasets, UGC-VideoCap emphasizes balanced integration of audio and visual modalities, featuring 1000 TikTok videos annotated through a structured three stage human-in-the-loop pipeline covering audio only, visual only, and joint audio visual semantics. The benchmark also includes 4000 carefully crafted QA pairs probing both unimodal and cross modal understanding. Alongside the dataset, we propose UGC-VideoCaptioner(3B), a 3B parameter captioning model distilled from Gemini 2.5 Flash. Using a novel two-stage training strategy supervised fine tuning followed by Group Relative Policy Optimization (GRPO), our approach enables efficient adaptation from limited data while maintaining competitive performance. Together, our benchmark and model offer a high-quality foundation and a data-efficient solution for advancing omnimodal video captioning in unconstrained real-world UGC settings.

视频字幕多模态TikTok轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。