首次尝试让音频大模型用思维链推理,提升复杂听觉任务理解能力。
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
- 引入思维链(CoT)增强音频大模型的推理能力。
- 简单中等任务准确率显著提升,难题上反而易混淆。
- 推理路径越长越准,适合复杂指令跟随与进阶推理场景。
大型音频语言模型(LALMs)在语音识别、音频描述等听觉感知任务中表现优异,但其解决复杂现实问题所需的推理能力仍待探索。本文首次将思维链(CoT)推理机制引入LALMs,评估其在声音、音乐和语音三大领域中的信息提取与推理性能。结果表明,CoT在简单与中等难度任务中显著提升准确率,但在高难度任务中推理链可能引发混淆。同时发现推理路径长度与准确率呈正相关,揭示了通过扩展推理规模来增强指令遵循与高级推理的潜力。本研究不仅验证了CoT在提升LALM推理能力上的前景,也指出了关键瓶颈,并为未来工作提供可行方向。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs) have demonstrated remarkable performance in tasks involving audio perception and understanding, such as speech recognition and audio captioning. However, their reasoning capabilities - critical for solving complex real-world problems - remain underexplored. In this work, we conduct the first exploration into integrating Chain-of-Thought (CoT) reasoning into LALMs to enhance their reasoning ability across auditory modalities. We evaluate representative CoT methods, analyzing their performance in both information extraction and reasoning tasks across sound, music, and speech domains. Our findings reveal that CoT methods significantly improve performance on easy and medium tasks but encounter challenges with hard tasks, where reasoning chains can confuse the model rather than improve accuracy. Additionally, we identify a positive correlation between reasoning path length and accuracy, demonstrating the potential of scaling inference for advanced instruction-following and reasoning. This study not only highlights the promise of CoT in enhancing LALM reasoning capabilities but also identifies key limitations and provides actionable directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。