arXiv:2604.21766cs.CL2026-04ACL

构建真实音频问答数据集,检验人类与AI的深层听觉推理能力

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

  • 设计基于真实音频的趣味问答题,需跨时序理解而非仅识别声音
  • 人类平均准确率32.13%,顶尖模型不足8.86%,差距显著
  • 采用项目反应理论分析模型缺陷,适合评估音频理解鲁棒性

现有音频问答基准多侧重声事件分类或基于字幕的提问,易导致模型通过捷径策略、短时线索、词汇先验或数据集特定偏见成功,甚至绕过音频依赖元数据和字幕。为此,我们提出AUDITA(来自多样化互联网趣闻作者的音频理解),一个大规模、真实世界的基准,用于严格评估超越表面声学识别的音频推理能力。AUDITA包含精心策划、由人类创作的趣闻问题,基于真实音频,通过具有挑战性的干扰项和长时程依赖关系,设计出无法仅凭孤立文本或声音线索回答的探测性问题。人类平均准确率为32.13%,既体现任务难度,也证明对音频有实质性理解。相比之下,当前最先进的音频问答模型表现极差,平均准确率低于8.86%。除原始准确率外,我们还应用项目反应理论(IRT)估计潜在能力、题目难度,并揭示模型与数据的系统性缺陷。

原文摘要 · Abstract (English)

Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed through shortcut strategies, short-duration cues, lexical priors, dataset-specific biases, or even bypassing audio via metadata and captions rather than genuine reasoning Thus, we present AUDITA (Audio Understanding from Diverse Internet Trivia Authors), a large-scale, real-world benchmark to rigorously evaluate audio reasoning beyond surface-level acoustic recognition. AUDITA comprises carefully curated, human-authored trivia questions grounded in real-world audio, designed to stress robust auditory reasoning through challenging distractors and long-range temporal dependencies, using probing queries that cannot be answered from isolated text or sound cues alone. Human average accuracy of 32.13% shows both the challenge of the task while demonstrating meaningful comprehension of the audio. In stark contrast, state of-the-art audio question answering models perform poorly, with average accuracy below 8.86%. Beyond raw accuracy, we apply Item Response Theory (IRT) to estimate latent proficiency, question difficulty, and expose systematic deficiencies of the models and data.

音频理解基准测试人类对比推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。