让大模型像人一样‘听懂’复杂声音,通过音频推理提升抗干扰能力。
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models

- 引入音视频思维链,让模型主动分析音频信号并进行实时处理。
- 在噪声环境下准确率下降超50%,而新方法使小模型提升24.73%、大模型达36.61%。
- 无需重训练即可增强鲁棒性,适合需要强抗干扰能力的语音系统。
近期的大规模音视频模型(LALMs)在语音翻译和音频问答等任务上表现强劲,但在复杂声学场景下的音频推理任务中仍存在显著局限。这类任务若能借助降噪、声源分离和精准时间对齐等声学工具将大有裨益,但现有LALMs缺乏对这些工具的访问能力。为此,我们提出Thinking-with-Sound(TwS)框架,通过结合语言推理与实时音频域分析,赋予LALMs音频思维链(Audio CoT)。不同于将音频视为静态输入的传统方法,TwS使模型能够主动“思考”音频信号,通过多模态推理执行数值分析与数字操作。为评估该方法,我们构建了MELD-Hard1k——一个通过引入多种声学扰动创建的新鲁棒性基准。实验表明,最先进LALMs在该基准上性能急剧下降,准确率相比干净音频降低超过50%。而TwS实现显著提升:小模型绝对准确率提高24.73%,大模型提升可达36.61%,且提升效果具可扩展性。研究证明,音频思维链可在不重新训练的前提下显著增强鲁棒性,为开发更稳健的音频理解系统开辟新路径。
原文摘要 · Abstract (English)
Recent Large Audio-Language Models (LALMs) have shown strong performance on various audio understanding tasks such as speech translation and Audio Q\&A. However, they exhibit significant limitations on challenging audio reasoning tasks in complex acoustic scenarios. These situations would greatly benefit from the use of acoustic tools like noise suppression, source separation, and precise temporal alignment, but current LALMs lack access to such tools. To address this limitation, we introduce Thinking-with-Sound (TwS), a framework that equips LALMs with Audio CoT by combining linguistic reasoning with on-the-fly audio-domain analysis. Unlike existing approaches that treat audio as static input, TwS enables models to actively think with audio signals, performing numerical analysis and digital manipulation through multimodal reasoning. To evaluate this approach, we construct MELD-Hard1k, a new robustness benchmark created by introducing various acoustic perturbations. Experiments reveal that state-of-the-art LALMs suffer dramatic performance degradation on MELD-Hard1k, with accuracy dropping by more than $50\%$ compared to clean audio. TwS achieves substantial improvements in robustness, demonstrating both effectiveness and scalability: small models gain $24.73\%$ absolute accuracy, with improvements scaling consistently up to $36.61\%$ for larger models. Our findings demonstrate that Audio CoT can significantly enhance robustness without retraining, opening new directions for developing more robust audio understanding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。