arXiv:2412.13543cs.CVcs.AI2024-12AAAI被引 2

提出新型音视频认知网络,提升视频片段检索与描述的准确性。

Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

  • 以查询为中心设计多模态感知机制,融合音视频全局对齐与局部交互。
  • 在HIREST数据集上达到当前最优性能,三项任务均显著领先。
  • 适用于需要精准理解用户查询的视频分析场景,如智能剪辑与摘要。

视频已成为互联网主流多媒体格式。为更好理解视频内容,新任务HIREST应运而生,涵盖视频检索、片段检索、片段分割和步骤描述。现有工作采用预训练CLIP模型进行视频检索,并作为特征提取器用于其余三项挑战性任务,采用多任务学习框架。然而,该方法因忽略模态间层次结构与关联关系,难以学习用户偏好内容的全面认知。本文基于由浅入深原则,提出查询中心音视频认知(QUAG)网络,构建可靠的多模态表示,用于片段检索、分割与步骤描述。具体地,首先设计模态协同感知模块,通过建模视觉与音频模态间的全局对比对齐与局部细粒度交互,获取丰富音视频内容;其次提出查询中心认知模块,利用深层查询对浅层音视频表示进行时序-通道过滤,从而认知用户偏好内容,生成面向查询的音视频表示。大量实验表明,QUAG在HIREST基准上取得最先进结果。进一步测试显示,其在基于查询的视频摘要任务中也表现出良好泛化能力。

原文摘要 · Abstract (English)

Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval, and leverages it as a feature extractor for other three challenging tasks solved in a multi-task learning paradigm. Nevertheless, this work struggles to learn the comprehensive cognition of user-preferred content, due to disregarding the hierarchies and association relations across modalities. In this paper, guided by the shallow-to-deep principle, we propose a query-centric audio-visual cognition (QUAG) network to construct a reliable multi-modal representation for moment retrieval, segmentation and step-captioning. Specifically, we first design the modality-synergistic perception to obtain rich audio-visual content, by modeling global contrastive alignment and local fine-grained interaction between visual and audio modalities. Then, we devise the query-centric cognition that uses the deep-level query to perform the temporal-channel filtration on the shallow-level audio-visual representation. This can cognize user-preferred content and thus attain a query-centric audio-visual representation for three tasks. Extensive experiments show QUAG achieves the SOTA results on HIREST. Further, we test QUAG on the query-based video summarization task and verify its good generalization.

视频理解多模态查询感知任务融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。