专注视频语音内容,提升图文检索准确率
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
- 增设语音专用分支,更好提取语音语义
- 早期视觉音频对齐,融合效果更优
- 在多个数据集上显著超越现有方法
在视频-文本检索任务中,CLIP已成为主流选择。由于其仅提供图像和文本编码器,导致现有方法完全忽略视频音轨。尽管已有研究尝试引入音频,通常通过添加音频编码器并融合视觉特征,但仍面临语音内容表征不足与视觉-音频融合不佳的问题。为此,我们提出SAVE(Speech-Aware Video Representation Learning)方法,改进当前最优的AVIGATE模型:设计专用语音分支以更有效生成语音嵌入,并引入soft-ALBEF实现早期视觉-音频对齐,促进融合。在五个基准数据集上的大量实验表明,SAVE性能显著优于现有方法,在SumR指标下,分别比AVIGATE提升4.1%(MSRVTT-9k)、1.9%(MSRVTT-7k)、2.5%(VATEX)、9.8%(Charades)和2.1%(LSMDC)。
原文摘要 · Abstract (English)
For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have been made to reintroduce audio -- typically by incorporating an audio encoder and fusing its output with visual features -- these methods face two challenges: ineffective representation of speech content and suboptimal vision-audio fusion. To address these issues jointly, we propose SAVE, a Speech Aware Video rEpresentation learning method. SAVE improves upon AVIGATE, a SOTA audiovisual method, with a dedicated speech branch for more effective speech embedding. Furthermore, we introduce soft-ALBEF for early vision-audio alignment that facilitates fusion. Extensive experiments on five benchmarks show that SAVE compares favorably against the SOTA, outperforming AVIGATE by +4.1% on MSRVTT-9k, +1.9% on MSRVTT-7k, +2.5% on VATEX, +9.8% on Charades, and +2.1% on LSMDC, in light of the SumR metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。