构建音频组合推理新基准,测试模型对多重声音事件的分析能力
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
- 设计五类任务覆盖计数、分类、检测等复合推理场景
- 实测主流大音频模型在复杂声音环境下降80%以上准确率
- 适合研究音频理解、多事件建模与模型鲁棒性的学者
大型音频语言模型(LALMs)在音频推理方面能力日益增强,但现有基准对多重声音共现场景下的组合推理覆盖有限。为填补这一空白,我们提出PolyBench,一个专用于评估多声部音频中组合推理能力的基准,包含五个评估子集:计数、分类、检测、并发性判断和时长估计,均需对多个共现事件及其关系进行推理。对当前顶尖LALMs的评估显示,其在多声部场景下性能普遍下降超过80%,表明现有模型在处理复杂音频结构方面存在根本性瓶颈。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio, yet existing benchmarks offer limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional structure. To address this gap, we introduce PolyBench, a benchmark designed to evaluate compositional reasoning in polyphonic audio, comprising five evaluation subsets that cover counting, classification, detection, concurrency, and duration estimation, all of which require reasoning over multiple concurrent events and their relations. Our evaluation of state-of-the-art LALMs reveals consistent performance degradation in polyphonic settings, indicating a fundamental bottleneck in current LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。