构建多领域音频问答基准,测试模型对复杂声音场景的理解与推理能力。
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
- 设计三个子任务:生物声学、时间声景和复杂问题,覆盖多样声音场景。
- 基于开发集结果,不同模型在各子任务上表现差异显著,最高准确率达72.3%。
- 适合研究音频语言模型、跨模态推理及智能体听觉感知的学者与工程师。
我们介绍 DCASE 2025 挑战赛的任务5:一个涵盖多个声音理解领域的音频问答(AQA)基准。该任务定义了三个问答子集(生物声学、时间声景、复杂问答),用于测试音频-语言模型在多样化声学场景下的交互式问答能力。数据集包含从海洋哺乳动物叫声到自然声景及真实世界复杂片段的多源音频。评估协议采用顶1准确率,并引入答案打乱鲁棒性检验。基线系统包括 Qwen2-Audio-7B、AudioFlamingo 2 与 Gemini-2-Flash。在开发集上的初步结果表明,不同模型在各子任务间表现差异明显,最高准确率为72.3%。本挑战旨在推动音频-语言模型向人类水平的听觉敏锐度演进,这对实现智能体有效感知与互动世界至关重要。
原文摘要 · Abstract (English)
We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over diverse acoustic scenes. We describe the dataset composition (from marine mammal calls to soundscapes and complex real-world clips), the evaluation protocol (top-1 accuracy with answer-shuffling robustness), and baseline systems (Qwen2-Audio-7B, AudioFlamingo 2, Gemini-2-Flash). Preliminary results on the development set are compared, showing strong variation across models and subsets. This challenge aims to advance the audio understanding and reasoning capabilities of audio-language models toward human-level acuity, which are crucial for enabling AI agents to perceive and interact about the world effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。