构建音频理解新基准,覆盖无障碍与工业场景真实需求
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
- 设计四类真实场景音频任务,涵盖背景音、噪声定位等非语音理解
- 测试五款顶尖大模型,发现部分任务准确率低于随机猜测
- 兼顾性能与延迟,验证合成数据生态效度,指导模型优化方向
大语言模型推动音频处理进步,催生了大型音频语言模型(LALMs)。然而,当前对语音识别之外的音频理解评估仍严重不足。本文提出SCENEBench基准套件,聚焦四大现实应用场景:背景声音理解、噪声定位、跨语言语音理解及发声者特征识别,这些均源于辅助技术和工业噪声监控中的未充分研究需求。该基准不仅评估模型性能,还测量延迟。所有音频样本通过合成方式生成(如叠加两个自然音频),并进一步用20个每任务的真实音频项(从现有数据集中筛选匹配任务条件)验证生态有效性。评估五款前沿LALMs后发现,模型表现差异显著,部分任务准确率甚至低于随机水平,另一些则达高精度。结果揭示了模型能力的关键短板,为针对性改进提供依据。
原文摘要 · Abstract (English)
Advances in large language models (LLMs) have enabled significant capabilities in audio processing, resulting in state-of-the-art models now known as Large Audio Language Models (LALMs). However, minimal work has been done to measure audio understanding beyond automatic speech recognition (ASR). This paper closes that gap by proposing a benchmark suite, SCENEBench (Spatial, Cross-lingual, Environmental, Non-speech Evaluation), that targets a broad form of audio comprehension across four real-world categories: background sound understanding, noise localization, cross-linguistic speech understanding, and vocal characterizer recognition. These four categories are selected based on understudied needs from accessibility technology and industrial noise monitoring. In addition to performance, we also measure model latency. The purpose of this benchmark suite is to assess audio beyond just what words are said - rather, how they are said and the non-speech components of the audio. Because our audio samples are synthetically constructed (e.g., by overlaying two natural audio samples), we further validate our benchmark against 20 natural audio items per task, sub-sampled from existing datasets to match our task criteria, to assess ecological validity. We assess five state-of-the-art LALMs and find critical gaps: performance varies across tasks, with some tasks performing below random chance and others achieving high accuracy. These results provide direction for targeted improvements in model capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。