提出分层对齐与解耦表征,提升音视频文本检索精度
Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
- 用双流Transformer实现多层级音文对齐
- 解耦特征捕捉细粒度语义关联,性能显著提升
- 自适应融合机制增强局部语义匹配能力
现有音频-文本检索方法多依赖单层次交互,难以充分对齐不同模态,导致匹配效果不佳。本文提出一种新框架,结合双流Transformer与分层对齐(THA)模块,识别音频与文本在不同Transformer层间的多层级对应关系。同时,针对现有方法忽视细粒度语义的缺陷,引入解耦跨模态表示(DCR)方法,将高维特征分解为紧凑潜在因子,以捕捉更精细的音文语义关联。此外,设计置信度感知(CA)模块,评估每对潜在因子的置信度,并自适应聚合跨模态因子,实现局部语义对齐。实验表明,THA有效提升检索性能,DCR进一步带来稳定增益。
原文摘要 · Abstract (English)
Most existing audio-text retrieval (ATR) approaches typically rely on a single-level interaction to associate audio and text, limiting their ability to align different modalities and leading to suboptimal matches. In this work, we present a novel ATR framework that leverages two-stream Transformers in conjunction with a Hierarchical Alignment (THA) module to identify multi-level correspondences of different Transformer blocks between audio and text. Moreover, current ATR methods mainly focus on learning a global-level representation, missing out on intricate details to capture audio occurrences that correspond to textual semantics. To bridge this gap, we introduce a Disentangled Cross-modal Representation (DCR) approach that disentangles high-dimensional features into compact latent factors to grasp fine-grained audio-text semantic correlations. Additionally, we develop a confidence-aware (CA) module to estimate the confidence of each latent factor pair and adaptively aggregate cross-modal latent factors to achieve local semantic alignment. Experiments show that our THA effectively boosts ATR performance, with the DCR approach further contributing to consistent performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。