arXiv:2603.29263cs.SD2026-03

测试大模型是否真听懂音频,发现极易被误导产生幻觉。

Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models

  • 设计6500个问答对,通过提问和伪造音频诱导模型编造不存在的声音。
  • 顶尖模型攻击成功率高达95.35%,表明现有评测无法反映真实可靠性。
  • 提出新数据集AHA-Guard,可降低49%的幻觉攻击成功率,适合安全敏感场景使用。

大型音频语言模型(LALMs)在音频-语言任务中表现优异,但其在真实场景中的可靠性尚未充分探索。本文提出音频幻觉攻击(AHA),构建包含6500个问答对的AHA-Eval评估套件,用于检验LALMs是否真正基于音频输入生成回答。该攻击针对两个方向:(i) 基于问题结构的攻击,诱导模型对未出现的声音产生幻觉;(ii) 基于音频的攻击,向音频流注入合成语音描述不存在的事件。在评估Audio Flamingo 3与Gemini 3 Pro等先进模型时,观察到高达95.35%和79.65%的攻击成功率达,暴露出标准基准性能掩盖的可靠性差距。为缓解此问题,我们提出一个12万条问答的后对齐数据集AHA-Guard,可将攻击成功率最高降低49%。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) achieve strong performance on audio-language tasks; however, their reliability in real-world settings remains underexplored. We introduce Audio Hallucination Attacks (AHA), an attack suite called AHA-Eval, comprising 6.5K QA pairs designed to test whether LALMs genuinely ground their responses in the audio input. AHA targets two attack surfaces: (i) query-based attacks, which exploit question structure to induce hallucinations about absent sounds, and (ii) audio-based attacks, which inject synthetic speech describing non-existent events into the audio stream. Evaluating state-of-the-art LALMs, including Audio Flamingo 3 and Gemini 3 Pro, we observe high attack success rates of 95.35% and 79.65%, respectively, revealing a reliability gap that is hidden by standard benchmark performance. To mitigate this, we propose a 120K QA post-alignment dataset, AHA-Guard, which successfully reduces attack success rates by up to 49%.

音频模型幻觉检测安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。