通过门控注意力机制精准利用音频信息,提升视频文本检索效果
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
- 用门控注意力筛选有用音频,避免无关音效干扰
- 在多个公开数据集上达到当前最优性能
- 适合需要融合多模态信息的视频理解任务
视频-文本检索任务旨在根据文本查询检索视频或反之,对视频理解与多模态信息检索至关重要。现有方法主要依赖视觉与文本特征,常忽略音频信息,尽管音频有助于整体内容理解。传统模型盲目使用音频输入,无论其是否有效,导致视频表征不佳。为此,本文提出一种新框架AVIGATE(Audio-guided VIdeo representation learning with GATEd attention),通过门控注意力机制有选择地利用音频线索,过滤无效音频信号。同时,设计自适应边距对比损失,缓解视频与文本间正负样本关系模糊的问题,促进更优对齐。大量实验表明,AVIGATE在所有公开基准上均取得领先性能。
原文摘要 · Abstract (English)
Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and textual features and often ignore audio, although it helps enhance overall comprehension of video content. Moreover, traditional models that incorporate audio blindly utilize the audio input regardless of whether it is useful or not, resulting in suboptimal video representation. To address these limitations, we propose a novel video-text retrieval framework, Audio-guided VIdeo representation learning with GATEd attention (AVIGATE), that effectively leverages audio cues through a gated attention mechanism that selectively filters out uninformative audio signals. In addition, we propose an adaptive margin-based contrastive loss to deal with the inherently unclear positive-negative relationship between video and text, which facilitates learning better video-text alignment. Our extensive experiments demonstrate that AVIGATE achieves state-of-the-art performance on all the public benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。