构建新基准,评估大模型在复杂声音场景中的上下文理解能力。
From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

- 通过真实与合成音源组合生成时序精准的半合成音频流。
- 多模型测试显示需融合语音、声效与环境音才能准确理解场景。
- 适合关注音频理解、多模态推理的研究者和开发者。
近期大型音频语言模型(LALMs)在语音、声响和音乐等单一声学层面上取得了显著进展。然而,现有评测基准大多孤立评估各层面,忽略了真实听觉场景中多种声源共存时的复杂上下文关系。真实世界的声音理解需要上下文感知的听觉场景理解(CASU),即通过整合语音、声学事件(如广播)和背景环境(如交通)等多层次信息来把握整体场景。为此,我们提出CASU基准,评估模型是否能理解由语音、声学事件和背景环境构成的听觉场景,并推理各层间的逻辑关系。我们设计了一套可扩展的流水线,通过组合真实场景音与合成语音生成时间精准的半合成音频流。基于此数据,构建了四项任务:情境问答、场景实体提取、说话人角色推断及反事实推理(场景被修改)。多个LALMs的实验表明,有效的听觉场景理解必须整合所有声学层,而非仅依赖语音或声响,凸显了实现复杂音频理解对CASU的必要性。
原文摘要 · Abstract (English)
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes. Real-world auditory interpretation requires Context-Aware Auditory Scene Understanding (CASU): the ability to comprehend the holistic scene by integrating sound layers. To evaluate this capability, we introduce the CASU benchmark, which assesses whether Audio LLMs can interpret auditory scenes composed of speech, acoustic events (e.g., announcements), and background environments (e.g., traffic), and reason about the logical relationships between these layers. We propose a scalable pipeline for constructing time-accurate, semi-synthetic audio streams by composing real-world scene sounds with synthetic speech. Building on this data, we design four tasks that probe scene understanding: contextual question answering, entity extraction from the scene, speaker role inference, and counterfactual reasoning where scene is manipulated. Experiments across multiple LALMs demonstrate that effective auditory scene understanding requires integration over all auditory layers, rather than reliance on speech or sound alone, underscoring the necessity of CASU for advancing complex audio understanding in LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。