arXiv:2505.03319cs.CVcs.AI2025-05被引 4

根据用户脚本生成精准视频摘要,实现内容定制化。

SD-VSum: A Method and Dataset for Script-Driven Video Summarization

  • 用跨模态注意力对齐文本与视频内容,动态选择相关片段。
  • 在VideoXum数据集上构建脚本描述,支持多版本摘要生成。
  • 可按用户需求生成不同侧重的摘要,适合个性化推荐场景。

本文提出脚本驱动的视频摘要任务,即根据用户提供的脚本(描述期望摘要的视觉内容)来选取视频中相关片段生成摘要。为此,我们扩展了现有的大规模通用视频摘要数据集VideoXum,为每个人工标注的摘要添加自然语言描述,形成“视频-摘要-摘要描述”三元组,使数据集适配新任务。基于此,我们设计了SD-VSum网络架构,采用跨模态注意力机制融合视觉与文本信息。实验表明,该方法在查询驱动及通用(单模态、多模态)摘要任务中均优于现有最先进方法,能够根据用户对内容的偏好生成定制化视频摘要。

原文摘要 · Abstract (English)

In this work, we introduce the task of script-driven video summarization, which aims to produce a summary of the full-length video by selecting the parts that are most relevant to a user-provided script outlining the visual content of the desired summary. Following, we extend a recently-introduced large-scale dataset for generic video summarization (VideoXum) by producing natural language descriptions of the different human-annotated summaries that are available per video. In this way we make it compatible with the introduced task, since the available triplets of ``video, summary and summary description'' can be used for training a method that is able to produce different summaries for a given video, driven by the provided script about the content of each summary. Finally, we develop a new network architecture for script-driven video summarization (SD-VSum), that employs a cross-modal attention mechanism for aligning and fusing information from the visual and text modalities. Our experimental evaluations demonstrate the advanced performance of SD-VSum against SOTA approaches for query-driven and generic (unimodal and multimodal) summarization from the literature, and document its capacity to produce video summaries that are adapted to each user's needs about their content.

视频摘要跨模态脚本驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。