首篇系统综述音频推理,厘清模型设计与应用范式。
A Survey of Audio Reasoning in Multimodal Foundation Models

- 区分直接预测与推理增强生成两类范式
- 梳理音频模态下多任务推理进展与挑战
- 适合关注语音理解、多模态系统的研究者
推理已成为现代基础模型的核心能力,但音频模态的推理发展仍受限。音频具有连续性、时间密集性,并在多时间尺度上包含语言、副语言和环境信息,其信号需与大语言模型的离散语义空间对齐,同时保留细粒度信息以实现可靠推理。当前进展受三大障碍制约:真实音频接地推理数据稀缺、捷径学习与模态幻觉问题,以及推理深度与语音交互实时延迟之间的矛盾。本文首次系统综述音频推理,提出统一范式,区分直接预测建模与推理增强生成;回顾音频推理模型的架构与训练基础;系统整理音频到文本、音频到语音、音视频推理及代理式音频推理的最新进展。进一步探讨链式思维提示、监督微调、强化学习及低延迟语音交互等新兴范式,分析评估方法、开放挑战与未来方向。旨在为构建鲁棒、高效且原生接地的音频推理系统提供清晰路线图。
原文摘要 · Abstract (English)
Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from those of text and vision. It is continuous, temporally dense, and contains linguistic, paralinguistic, and environmental information at multiple time scales. As a result, audio reasoning models must align acoustic signals with the discrete semantic space of large language models, while still preserving fine-grained information needed for reliable inference. Progress is also limited by three major obstacles: the scarcity of genuinely audio-grounded reasoning data, shortcut learning and modality hallucination, and the tension between reasoning depth and real-time latency in spoken interaction. In this paper, we present the first dedicated survey of audio reasoning. We provide a unified formulation that distinguishes direct predictive modeling from reasoning-augmented generation, review the architectural and training foundations of audio reasoning models, and systematically organize recent advances in Audio-to-Text, Audio-to-Speech, Audio-Visual Reasoning and Agentic Audio Reasoning. We further examine emerging paradigms such as Chain-of-Thought prompting, supervised fine-tuning, reinforcement learning, and latency-aware spoken interaction, and discuss evaluation practices, open challenges, and future directions. Our goal is to offer a coherent roadmap for developing robust, efficient, and natively grounded audio reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。