MUGEN benchmark揭示大音频模型在多音频理解中的瓶颈与改进方法
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
- 构建多音频理解评估基准,覆盖语音、通用音频和音乐三类场景
- 输入音频数量增加时性能显著下降,暴露缩放瓶颈问题
- 通过音频顺序打乱提升预测鲁棒性,最高增益6.28%,适合音频研究者
尽管多音频理解对大型音频语言模型(LALMs)至关重要,但该领域仍缺乏深入探索。本文提出 MUGEN,一个全面的基准测试,涵盖语音、通用音频和音乐三类任务。实验发现,在多音频设置下模型表现持续薄弱,且随着并发音频输入数量增加,性能急剧下降,揭示了输入缩放是根本性瓶颈。我们进一步考察无需训练的策略,发现音频顺序自一致性方法(Audio-Permutational Self-Consistency)通过多样化音频候选顺序,能帮助模型形成更稳健的聚合预测,带来最高达6.28%的准确率提升。结合思维链(Chain-of-Thought)后性能进一步提升至6.74%。这些结果揭示了当前LALMs在复杂听觉理解中的盲点,并为未来评估提供了基础。
原文摘要 · Abstract (English)
While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-of-Thought further improves performance to 6.74%. These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。