arXiv:2510.05652cs.CV2025-10

基于脚本的多模态视频摘要方法,提升与语音内容的相关性

SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets

  • 引入加权跨模态注意力,融合脚本与视频、语音文本的语义相关性
  • 在两个扩展数据集上表现优于现有最先进方法,尤其在脚本匹配度上提升显著
  • 适合需要精准脚本驱动视频压缩的研究者或应用开发人员

本文提出一种脚本驱动的多模态视频摘要方法SD-MVSum及两个大规模数据集。该方法在先前仅考虑视觉内容的SD-VSum基础上,新增对用户脚本与视频语音转录文本之间相关性的建模。通过一种新的加权跨模态注意力机制,显式利用脚本与视频、脚本与转录文本之间的语义相似性,强化与脚本最相关的视频片段。同时,我们扩展了S-VideoXum和MrHiSum两个大规模数据集,使其适用于脚本驱动的多模态视频摘要训练与评估。实验表明,SD-MVSum在脚本驱动和通用视频摘要任务中均达到当前最优性能。相关代码与数据集已开源:https://github.com/IDT-ITI/SD-MVSum。

原文摘要 · Abstract (English)

In this work, we present a method and two large-scale datasets for Script-Driven Multimodal Video Summarization. The proposed method, SD-MVSum, builds on our earlier SD-VSum method for script-driven video summarization, which considered just the visual content of the video. SD-MVSum takes into account, in addition to the visual modality, the relevance of the user-provided script with the spoken content (i.e., audio transcript) of the video. The dependence between each considered pair of data modalities, i.e., script-video and script-transcript, is modeled using a new weighted cross-modal attention mechanism. This mechanism explicitly exploits the semantic similarity between the paired modalities in order to promote the parts of the full-length video with the highest relevance to the user-provided script. Furthermore, we extend two large-scale datasets for script-driven (S-VideoXum) and generic (MrHiSum) video summarization, to make them suitable for training and evaluation of script-driven multimodal video summarization methods. Experimental comparisons document the competitiveness of the proposed SD-MVSum method against other SotA approaches for script-driven and generic video summarization. Our new method and extended datasets are available at: https://github.com/IDT-ITI/SD-MVSum.

视频摘要多模态脚本驱动注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。