首个百万级长视频章节模型,让小时视频可导航、可摘要。
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
- 构建跨语言多层级标注数据集,融合语音、文字与视觉信息。
- 在长视频章节划分上提升14.0% F1和11.3% SODA,刷新性能纪录。
- 支持迁移应用,显著改进下游视频密集描述任务效果。
小时级视频(如课程、播客、纪录片)的激增催生了高效内容结构化需求。然而现有方法受限于小规模标注数据,且标签短而粗糙,难以捕捉长视频中的细微转换。我们提出ARC-Chapter,首个基于百万级长视频章节的大规模视频章节模型,采用中英双语、时间对齐、分层标注。通过整合ASR转录、场景文本与视觉描述的结构化流程,构建从标题到长摘要的多层级标注数据集。实验表明,随着数据量与标注密度增加,性能显著提升。我们设计新评估指标GRACE,融合多对一片段重叠与语义相似性,更贴合真实章节划分灵活性。大量实验显示,ARC-Chapter以显著优势超越前序最佳模型:F1提升14.0%,SODA提升11.3%。同时具备优异迁移能力,在YouCook2数据集上的密集视频描述任务中也实现性能突破。
原文摘要 · Abstract (English)
The proliferation of hour-long videos (e.g., lectures, podcasts, documentaries) has intensified demand for efficient content structuring. However, existing approaches are constrained by small-scale training with annotations that are typical short and coarse, restricting generalization to nuanced transitions in long videos. We introduce ARC-Chapter, the first large-scale video chaptering model trained on over million-level long video chapters, featuring bilingual, temporally grounded, and hierarchical chapter annotations. To achieve this goal, we curated a bilingual English-Chinese chapter dataset via a structured pipeline that unifies ASR transcripts, scene texts, visual captions into multi-level annotations, from short title to long summaries. We demonstrate clear performance improvements with data scaling, both in data volume and label intensity. Moreover, we design a new evaluation metric termed GRACE, which incorporates many-to-one segment overlaps and semantic similarity, better reflecting real-world chaptering flexibility. Extensive experiments demonstrate that ARC-Chapter establishes a new state-of-the-art by a significant margin, outperforming the previous best by 14.0% in F1 score and 11.3% in SODA score. Moreover, ARC-Chapter shows excellent transferability, improving the state-of-the-art on downstream tasks like dense video captioning on YouCook2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。