首个面向语音、环境音与音乐的幻觉检测基准,揭示大音频模型严重缺陷。
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

- 构建5000+人工验证问答对,覆盖三类音频与四类任务
- 发现模型在声学依据、时间推理与音乐属性理解上普遍出错
- 支持细粒度分析,适合评估和改进音频模型可靠性
大型音频语言模型(LALMs)在各类音频任务中表现强劲,但其生成内容存在语义错误或声学不支持的幻觉问题仍缺乏系统研究。现有幻觉评测多集中于文本或视觉领域,少数音频相关研究规模小、模态覆盖窄且诊断深度不足。为此,我们提出HalluAudio,首个大规模跨语音、环境音与音乐的幻觉检测基准。该数据集包含超过5000个经人工验证的问答对,涵盖二元判断、多选推理、属性验证与开放式问答等多种任务类型。为系统诱发幻觉,我们设计对抗性提示与混合音频场景。评测协议不仅评估准确率,还量化幻觉率、是/否偏差、错误类型分析与拒答率,实现对模型失效模式的精细剖析。我们对多种开源及专有模型进行了全面评估,首次完成跨语音、声音与音乐的大规模对比。结果揭示模型在声学基础、时间推理与音乐属性理解方面存在显著不足,凸显构建可靠、鲁棒的音频语言模型的迫切需求。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain. Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth. We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music. HalluAudio comprises over 5K human-verified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA. To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions. Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes. We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music. Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。