动态融合视音频文三模态,提升长视频摘要精度
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
- 帧级自适应加权融合视觉、音频、文本信息
- 在4个数据集上超越现有方法,尤其在新基准上表现突出
- 首次发布含三模态的大型视频摘要数据集MoSu
视频内容的爆炸式增长亟需高效视频摘要技术以提取关键信息。然而,现有方法难以充分理解复杂视频,主要因其采用静态或模态无关的融合策略,无法捕捉视频中模态显著性随帧变化的动态特性。为此,我们提出TripleSumm,一种在帧级别自适应加权并融合视觉、文本和音频三模态贡献的新架构。此外,多模态视频摘要研究的一大瓶颈是缺乏全面的基准。为此,我们引入MoSu(Most Replayed Multimodal Video Summarization),首个提供三模态数据的大规模基准。大量实验表明,TripleSumm在包括MoSu在内的四个基准上均达到最优性能,显著优于现有方法。代码与数据集已开源。
原文摘要 · Abstract (English)
The exponential growth of video content necessitates effective video summarization to efficiently extract key information from long videos. However, current approaches struggle to fully comprehend complex videos, primarily because they employ static or modality-agnostic fusion strategies. These methods fail to account for the dynamic, frame-dependent variations in modality saliency inherent in video data. To overcome these limitations, we propose TripleSumm, a novel architecture that adaptively weights and fuses the contributions of visual, text, and audio modalities at the frame level. Furthermore, a significant bottleneck for research into multimodal video summarization has been the lack of comprehensive benchmarks. Addressing this bottleneck, we introduce MoSu (Most Replayed Multimodal Video Summarization), the first large-scale benchmark that provides all three modalities. Extensive experiments demonstrate that TripleSumm achieves state-of-the-art performance, outperforming existing methods by a significant margin on four benchmarks, including MoSu. Our code and dataset are available at https://github.com/smkim37/TripleSumm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。