arXiv:2604.23323cs.CLcs.SD2026-04

提升音频文本检索在噪声和长音频下的鲁棒性

Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss

论文配图:Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
图 1 · 摘自论文原文
  • 用跨模态注意力与混合损失优化音文嵌入
  • 小批次训练下仍保持稳定,长音频信噪比5~15仍有效
  • 适合嘈杂环境或长时音频的多媒体检索应用

音频文本检索实现音频内容与自然语言查询之间的语义对齐,支持多媒体搜索、无障碍访问和监控等应用。然而,当前先进方法因依赖对比学习和大批次训练,在处理长时、噪声大、弱标注的音频时表现不佳。本文提出一种新型多模态检索框架,通过结合Transformer投影、线性映射和双向注意力的跨模态嵌入精炼模块,优化音文嵌入。为进一步提升鲁棒性,引入融合余弦相似度、$´\mathcal{L}_{1}$ 和对比目标的混合损失函数,即使在小批次条件下也能实现稳定训练。该方法通过静音感知分块和基于注意力的池化,高效处理长音频与噪声数据(信噪比5至15)。在基准数据集上的实验表明,性能优于先前方法。

原文摘要 · Abstract (English)

Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. However, current state-of-the-art approaches struggle with long, noisy, and weakly labeled audio due to their reliance on contrastive learning and large-batch training. We propose a novel multimodal retrieval framework that refines audio and text embeddings using a cross-modal embedding refinement module combining transformer-based projection, linear mapping, and bidirectional attention. To further improve robustness, we introduce a hybrid loss function blending cosine similarity, $\mathcal{L}_{1}$, and contrastive objectives, enabling stable training even under small-batch constraints. Our approach efficiently handles long-form and noisy audio (SNR 5 to 15) via silence-aware chunking and attention-based pooling. Experiments on benchmark datasets demonstrate improvements over prior methods.

音频检索跨模态混合损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。