arXiv:2512.19703eess.AScs.IR2025-12

通过动态知识更新提升音频文本检索准确率

ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval

  • 引入多粒度知识注入打破梯度局部瓶颈
  • 动态更新知识库,解决模型与知识错配问题
  • 适配不同模型,适合音频文本检索研究者

当前音频-文本检索(ATR)主流方法采用双编码器架构,通过小批量对比学习优化。然而,仅依赖批次内样本进行优化会带来梯度局部性瓶颈(GLB),导致声学歧义难以解决,稀有长尾概念学习受限。尽管外部知识注入可突破该瓶颈,但常引发表征漂移不匹配(RDM)问题——静态知识库与动态演化的编码器逐渐失配,使知识引导退化为噪声。为此,我们提出自适应自提升知识框架(ASK)。ASK通过多粒度知识注入打破GLB,采用动态精炼策略同步知识库与模型以缓解RDM,同时设计自适应可靠性加权机制,基于跨模态一致性过滤检索噪声。在多个基准测试上的大量实验表明,ASK在不同骨干网络下均持续达到新最优性能。

原文摘要 · Abstract (English)

The dominant paradigm for Audio-Text Retrieval (ATR) relies on dual-encoder architectures optimized via mini-batch contrastive learning. However, restricting optimization to local in-batch samples creates a fundamental limitation we term the Gradient Locality Bottleneck (GLB), which prevents the resolution of acoustic ambiguities and hinders the learning of rare long-tail concepts. While external knowledge injection can break this bottleneck, it often triggers a problem called Representation-Drift Mismatch (RDM), where a static knowledge base becomes misaligned with evolving encoders, degrading guidance into noise. To address these intertwined challenges, we propose the Adaptive Self-improving Knowledge (ASK) framework. ASK breaks the GLB via multi-grained knowledge injection and mitigates RDM through a dynamic refinement strategy that synchronizes the knowledge base with the model. Additionally, an adaptive reliability weighting scheme is employed to filter retrieval noise based on cross-modal consistency. Extensive experiments across multiple benchmarks demonstrate that ASK consistently achieves new state-of-the-art performance across various backbones.

音频检索知识增强跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。