通过音视频分帧门控融合与语义感知扰动,提升文本-视频检索精度
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
- 分帧门控融合模块按文本指导选择关键音视频帧
- 语义方差缩放扰动使文本嵌入更稳定且区分度更高
- 在多个数据集上超越主流方法,计算开销低
文本到视频检索需要语言与时间丰富的音视频信号间精确对齐。现有方法常侧重视觉线索而低估音频语义或依赖粗粒度融合策略,导致多模态表征不佳。本文提出GAIS框架,从表征与正则化双角度强化多模态对齐。首先,帧级门控融合(FGF)模块在文本引导下自适应整合音视频特征,实现细粒度的时间选择性。其次,语义方差缩放扰动(SVSP)机制以语义感知方式控制扰动幅度,正则化文本嵌入空间。二者互补:FGF通过选择性融合缩小模态差距,SVSP提升嵌入稳定性与判别性。在MSR-VTT、DiDeMo、LSMDC和VATEX上的大量实验表明,GAIS在多种检索指标上持续优于强基线,同时保持显著计算效率。
原文摘要 · Abstract (English)
Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutilizing audio semantics or relying on coarse fusion strategies, resulting in suboptimal multimodal representations. We introduce GAIS, a retrieval framework that strengthens multimodal alignment from both representation and regularization perspectives. First, a Frame-level Gated Fusion (FGF) module adaptively integrates audio-visual features under textual guidance, enabling fine-grained temporal selection of informative frames. Second, a Semantic Variance-Scaled Perturbation (SVSP) mechanism regularizes the text embedding space by controlling perturbation magnitude in a semantics-aware manner. These two modules are complementary: FGF minimizes modality gaps through selective fusion, while SVSP improves embedding stability and discrimination. Extensive experiments on MSR-VTT, DiDeMo, LSMDC, and VATEX demonstrate that GAIS consistently outperforms strong baselines across multiple retrieval metrics while maintaining notable computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。