构建开源音频推理系统,提升模型理解复杂音频的能力。
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models

- 用自蒸馏策略微调模型,增强音频推理能力。
- 生成54.5万条高质量音频推理数据,覆盖复杂任务场景。
- 在多个评测中表现领先,适合音频智能研究者使用。
近期推理模型在文本和多模态领域取得显著进展,但音频推理仍较薄弱。目前仅有少数大型音频语言模型(LALMs)引入显式的思维链(CoT)推理,且能力不稳定、难以应对复杂任务。为填补这一空白,我们提出 Audio-Cogito,一个完全开源的深度音频推理解决方案。我们开发了 Cogito-pipe 工具用于高质量音频推理数据的构建,生成了 54.5 万条推理样本。基于该数据集,采用自蒸馏策略进行模型微调。在唯一评估思维链过程的 MMAR 基准测试中,我们的模型在开源模型中表现最佳,并在特定指标上达到或超越部分闭源模型水平。该方法在 Interspeech 2026 音频推理挑战赛中也位列顶尖梯队。
原文摘要 · Abstract (English)
Recent advances in reasoning models have driven significant progress in text and multimodal domains, yet audio reasoning remains relatively limited. Only a few Large Audio Language Models (LALMs) incorporate explicit Chain-of-Thought (CoT) reasoning, and their capabilities are often inconsistent and insufficient for complex tasks. To bridge this gap, we introduce Audio-Cogito, a fully open-source solution for deep audio reasoning. We develop Cogito-pipe for high-quality audio reasoning data curation, producing 545k reasoning samples. Based on this dataset, we adopt a self-distillation strategy for model fine-tuning. Experiments on the MMAR benchmark, the only audio benchmark evaluating the CoT process, show that our model achieves the best performance among open-source models and matches or surpasses certain closed-source models in specific metrics. Our approach also ranks among the top-tier systems in the Interspeech 2026 Audio Reasoning Challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。