arXiv:2606.09925cs.SD2026-06被引 1

构建音频推理过程错误检测基准,评估大模型推理质量。

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning

论文配图:AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning
图 1 · 摘自论文原文
  • 构建分步标注的音频推理错误数据集,支持细粒度错误类型识别。
  • 在6个模型上测试,验证模型对音频特定错误类型的诊断能力差异。
  • 适用于音频推理验证器、过程奖励模型等研究方向。

大型音频-语言模型(LALMs)越来越多地使用显式推理链来理解复杂音频,但推理质量的评估仍不充分。尽管文本和多模态领域的过程奖励模型(PRMs)已有进程级基准,但针对音频推理的类似评估仍较匮乏。本文提出AudioProcessBench,一个面向音频推理步骤级错误识别的综合性基准。该基准包含6个音频与全模态语言模型生成的多样化推理链,每条链被分割为离散推理步骤,并标注二值正确性及细粒度错误类型。基准在三种互补范式下评估模型:(1)步骤正确性识别;(2)基于错误类型的条件检测,用于诊断音频特定验证器能力;(3)链级聚合,验证器在相同问题的不同推理链间选择或融合最优答案。该设计使我们能系统分析当前模型是否能检测过程错误、其弱点是否因音频特有错误类型而异,以及过程验证是否提升最终答案选择。AudioProcessBench为未来音频推理验证器、过程奖励模型和可靠全模态推理研究提供测试平台。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) increasingly use explicit reasoning traces for complex audio understanding, yet the evaluation of reasoning quality remains underexplored. Although process-level benchmarks for process reward models (PRMs) have advanced reasoning evaluation in text and multi-modal domains, comparable evaluation for audio reasoning remains limited. In this paper, we present AudioProcessBench, a comprehensive benchmark for step-level process error identification in audio reasoning. AudioProcessBench contains diverse reasoning traces generated by 6 audio and omni language models. Each trace is segmented into discrete reasoning steps and annotated with binary step correctness and fine-grained error types. Our benchmark evaluates models under three complementary paradigms: (1) step correctness identification, (2) error-type-conditioned detection for diagnosing audio-specific verifier capacities, and (3) chain-level aggregation, where verifiers select or aggregate among multiple reasoning traces for the same question. This design enables a systematic analysis of whether current models can detect process errors, whether their weaknesses differ across audio-specific error types, and whether process verification translates into improved answer selection. AudioProcessBench provides a testbed for future research on audio reasoning verifiers, process reward models, and reliable omni-modal reasoning.

音频理解推理验证基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。