用鲸类叫声测试大模型听觉感知能力,发现差距显著。
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
- 用海洋哺乳动物叫声构建听觉评测基准
- 模型在音高/时长识别上远低于人类水平
- 适合研究音频语言模型听觉理解的学者
大型音频语言模型(LALMs)将语言理解拓展至听觉领域,但其对音高、时长等低层听觉感知能力仍缺乏深入评估。而此类能力对处理未知声音的现实任务至关重要。为此,我们提出世界鲸类评测基准(WoW-Bench),基于海洋哺乳动物叫声评估低层听觉感知与认知能力。该基准包含感知子任务(分类新声音)与认知子任务(参照布卢姆分类法,评估记忆、理解、应用、分析能力),并引入干扰问题以检验模型是否真正通过听觉推理而非其他启发式策略解题。对主流LALMs的实验显示,其表现远低于人类水平,表明当前模型仍需更强的听觉基础支撑。
原文摘要 · Abstract (English)
Large audio language models (LALMs) extend language understanding into the auditory domain, yet their ability to perform low-level listening, such as pitch and duration detection, remains underexplored. However, low-level listening is critical for real-world, out-of-distribution tasks where models must reason about unfamiliar sounds based on fine-grained acoustic cues. To address this gap, we introduce the World-of-Whale benchmark (WoW-Bench) to evaluate low-level auditory perception and cognition using marine mammal vocalizations. WoW-bench is composed of a Perception benchmark for categorizing novel sounds and a Cognition benchmark, inspired by Bloom's taxonomy, to assess the abilities to remember, understand, apply, and analyze sound events. For the Cognition benchmark, we additionally introduce distractor questions to evaluate whether models are truly solving problems through listening rather than relying on other heuristics. Experiments with state-of-the-art LALMs show performance far below human levels, indicating a need for stronger auditory grounding in LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。