arXiv:2502.05435eess.AScs.AI2025-02NeurIPS被引 1

提出新型核函数,解决音频描述生成中的时序偏差问题。

Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning

  • 用无偏切片Wasserstein核+旋转位置编码,保留音语模态时间关系
  • 在AudioCaps和Clotho上提升描述质量与文本多样性,检索准确率更高
  • 可通用至音频推理任务,使大模型推理准确率提升4%

音频描述系统面临根本挑战:教师强制训练会引发暴露偏差,导致推理时描述退化。尽管已有对比学习方法作为解决方案,但通常无法捕捉声学与语言模态间的关键时序关系。本文提出一种带旋转位置编码的无偏切片Wasserstein RBF(USW-RBF)核,专门用于保持跨模态的时间信息。该方法具有实际优势:核函数支持高效的随机梯度优化,适用于真实场景。基于此,我们构建了完整的音频描述框架,并引入随机解码以进一步缓解描述退化。在AudioCaps和Clotho数据集上的大量实验表明,该方法显著提升了描述质量、词汇多样性及文本到音频检索准确率。此外,我们将USW-RBF核扩展至音频推理任务,在CompA-R上提升了大型音频语言模型的推理正确性与质量;在MMAU-test-mini基准上,推理准确率提升4%。结果证明该方法是应对音频-语言任务中跨模态对齐挑战的强大且通用的解决方案。

原文摘要 · Abstract (English)

Audio captioning systems face a fundamental challenge: teacher-forcing training creates exposure bias that leads to caption degeneration during inference. While contrastive methods have been proposed as solutions, they typically fail to capture the crucial temporal relationships between acoustic and linguistic modalities. We address this limitation by introducing the unbiased sliced Wasserstein RBF (USW-RBF) kernel with rotary positional embedding, specifically designed to preserve temporal information across modalities. Our approach offers a practical advantage: the kernel enables efficient stochastic gradient optimization, making it computationally feasible for real-world applications. Building on this foundation, we develop a complete audio captioning framework that integrates stochastic decoding to further mitigate caption degeneration. Extensive experiments on AudioCaps and Clotho datasets demonstrate that our method significantly improves caption quality, lexical diversity, and text-to-audio retrieval accuracy. Furthermore, we demonstrate the generalizability of our USW-RBF kernel by applying it to audio reasoning tasks, where it enhances the reasoning capabilities of large audio language models on the CompA-R in terms of correctness and quality. Our kernel also improves the reasoning accuracy of the MMAU-test-mini benchmarks by $4\%$. These results establish our approach as a powerful and generalizable solution for cross-modal alignment challenges in audio-language tasks.

音频生成跨模态核方法时序对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。