解决音频模型推理中感知能力随步骤衰减的问题
When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
- 用多步感知引导推理,动态分解复杂问题
- 在CAFE上感知准确率从31.74%提升至63.51%
- 适合研究音频理解与长序列推理的学者
测试时扩展在通过增加推理计算量解决复杂问题方面表现出显著效果。然而,在大型音频-语言模型(LALMs)中存在一种反直觉现象:对结构化推理轨迹进行训练,相比直接回答,性能提升微弱甚至下降。为探究该现象,我们提出CAFE评估框架,可精确量化音频推理错误。评估结果表明,LALMs在推理过程中面临感知能力衰退的瓶颈,随着推理长度增加,感知性能显著下降。为此,我们提出MPAR²,一种鼓励动态感知推理并分解复杂问题为感知丰富的子任务的范式。借助强化学习,MPAR²将CAFE上的感知准确率从31.74%提升至63.51%,有效缓解感知衰减,并同步提升推理能力,在MMAU基准上达到74.59%的准确率。进一步分析表明,MPAR²增强了模型对音频输入的关注,并根据任务复杂度动态调整推理资源分配。
原文摘要 · Abstract (English)
Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive phenomenon exists: post-training models for structured reasoning trajectories results in marginal or even negative gains compared to post-training for direct answering. To investigate it, we introduce CAFE, an evaluation framework designed to precisely quantify audio reasoning errors. Evaluation results reveal LALMs struggle with perception during reasoning and encounter a critical bottleneck: reasoning performance suffers from audio perception decay as reasoning length extends. To address it, we propose MPAR$^2$, a paradigm that encourages dynamic perceptual reasoning and decomposes complex questions into perception-rich sub-problems. Leveraging reinforcement learning, MPAR$^2$ improves perception performance on CAFE from 31.74% to 63.51% and effectively mitigates perception decay, concurrently enhancing reasoning capabilities to achieve a significant 74.59% accuracy on the MMAU benchmark. Further analysis demonstrates that MPAR$^2$ reinforces LALMs to attend to audio input and dynamically adapts reasoning budget to match task complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。