arXiv:2410.14087cs.CV2024-10被引 4

根据用户提问生成更相关长视频摘要的新方法。

Your Interest, Your Summaries: Query-Focused Long Video Summarization

  • 用注意力和卷积网络提取查询相关的视频内容
  • 在基准数据集上显著提升摘要相关性
  • 适合需要精准视频摘要的用户

从长视频生成简洁且信息丰富的摘要具有重要意义,但其主观性强,因场景重要性差异而异。用户通过文本查询指定场景重要性可增强摘要的相关性。本文提出一种面向查询的视频摘要方法,旨在使摘要与用户查询高度对齐。为此,我们设计了全卷积序列网络带注意力机制(FCSNA-QFVS),利用时间卷积与注意力机制,有效提取并突出基于用户查询的相关内容。在面向查询视频摘要的基准数据集上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Generating a concise and informative video summary from a long video is important, yet subjective due to varying scene importance. Users' ability to specify scene importance through text queries enhances the relevance of such summaries. This paper introduces an approach for query-focused video summarization, aiming to align video summaries closely with user queries. To this end, we propose the Fully Convolutional Sequence Network with Attention (FCSNA-QFVS), a novel approach designed for this task. Leveraging temporal convolutional and attention mechanisms, our model effectively extracts and highlights relevant content based on user-specified queries. Experimental validation on a benchmark dataset for query-focused video summarization demonstrates the effectiveness of our approach.

视频摘要查询聚焦注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。