arXiv:2412.08879cs.CV2024-12被引 12

构建10万+片段数据集,推动用户视频长变短的自动转换研究

Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

  • 用两阶段标注法从真实用户内容中获取高质量剪辑标签
  • 收录超1万条视频、12万+片段,覆盖多样场景与风格
  • 融合音视频图文信息,提供跨模态复用基准模型

近年来,社交媒体平台对短时视频创作的需求显著增长。尽管视频摘要与精彩片段检测技术已有进展,但这些方法通常依赖特定领域知识,难以泛化至真实世界视频内容。为解决此问题,我们提出 Repurpose-10K,一个包含超过10,000条视频和120,000个标注片段的大规模数据集,旨在应对视频长变短的任务。针对非专业标注者可能带来的标注误差,我们设计了两阶段标注方案,基于真实用户生成内容获取更可靠的标签。此外,我们构建了一个基线模型,通过跨模态融合与对齐框架整合音频、视觉和字幕信息,以应对该挑战任务。本工作期望推动视频复用这一尚待深入探索领域的研究进展。

原文摘要 · Abstract (English)

The demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times. Despite notable advancements in the fields of video summarization and highlight detection, which can create partially usable short films from raw videos, these approaches are often domain-specific and require an in-depth understanding of real-world video content. To tackle this predicament, we propose Repurpose-10K, an extensive dataset comprising over 10,000 videos with more than 120,000 annotated clips aimed at resolving the video long-to-short task. Recognizing the inherent constraints posed by untrained human annotators, which can result in inaccurate annotations for repurposed videos, we propose a two-stage solution to obtain annotations from real-world user-generated content. Furthermore, we offer a baseline model to address this challenging task by integrating audio, visual, and caption aspects through a cross-modal fusion and alignment framework. We aspire for our work to ignite groundbreaking research in the lesser-explored realms of video repurposing.

视频复用多模态数据集用户生成内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。