arXiv:2511.13273cs.SDcs.AI2025-11被引 1

首个评测音频运动感知的基准,发现大模型在听觉空间推理上严重不足。

AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs

  • 构建首个针对声音运动感知的问答评测基准
  • 模型平均准确率低于50%,难以识别声音移动方向
  • 适合研究听觉认知与多模态模型的学者参考

大型音频语言模型(LALMs)在语音识别、音频描述和听觉问答方面取得了显著进展。然而,这些模型是否具备感知空间动态,特别是声源运动的能力仍不明确。本文揭示了当前LALMs存在系统性运动感知缺陷。为此,我们提出AudioMotionBench,首个专门用于评估音频语言模型听觉运动理解能力的基准。该基准采用受控的问答形式,考察模型能否从双耳音频中推断出移动声源的方向与轨迹。全面的定量与定性分析表明,现有模型难以可靠识别运动线索或区分方向模式,平均准确率低于50%,凸显听觉空间推理的根本局限。本研究揭示了人类与模型在听觉空间推理上的根本差距,为未来提升音频语言模型的空间认知能力提供了诊断工具与新洞见。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can perceive spatial dynamics, particularly the motion of sound sources, remains unclear. In this work, we uncover a systematic motion perception deficit in current ALLMs. To investigate this issue, we introduce AudioMotionBench, the first benchmark explicitly designed to evaluate auditory motion understanding. AudioMotionBench introduces a controlled question-answering benchmark designed to evaluate whether Audio-Language Models (LALMs) can infer the direction and trajectory of moving sound sources from binaural audio. Comprehensive quantitative and qualitative analyses reveal that current models struggle to reliably recognize motion cues or distinguish directional patterns. The average accuracy remains below 50\%, underscoring a fundamental limitation in auditory spatial reasoning. Our study highlights a fundamental gap between human and model auditory spatial reasoning, providing both a diagnostic tool and new insight for enhancing spatial cognition in future Audio-Language Models.

音频理解空间感知评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。