语音内容会诱导视听大模型幻觉,暴露其跨模态理解短板。
SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

- 构建首个语音-视觉幻觉评测基准SVHalluc,从语义与时间两维度诊断问题
- 开源模型在多任务中准确率接近随机,而Gemini 2.5 Pro表现显著更优
- 揭示当前模型虽单模态感知强,但缺乏语音引导的视频理解能力
尽管视听大语言模型(Audio-Visual LLMs)取得成功,但仍会产生看似合理却无依据的输出,即幻觉。现有基准主要关注环境音(如狗吠)来判断事件发生,但人类语音具有更丰富的语义和时间结构,当前模型是否能准确对齐语音内容与对应视觉信号仍未知。本工作表明,语音内容可诱发视听大模型幻觉。为此,我们提出SVHalluc——首个系统评估视听大模型语音-视觉幻觉的综合性基准。该基准从语义和时间两个关键且互补的角度诊断幻觉。实验结果表明,当前最先进的开源视听大模型在对齐语音内容与视觉信号方面表现不佳,多项任务准确率接近随机水平;相比之下,Gemini 2.5 Pro 显著优于开源模型。分析显示,其失败源于跨模态理解能力有限,尽管单模态感知能力强。本工作揭示了当前视听大模型的一项新且根本性的局限,并强调了语音引导视频理解的必要性。
原文摘要 · Abstract (English)
Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, rich semantics and temporal structures, yet it remains unexplored whether current models can accurately align speech content with corresponding visual signals. In this work, we show that speech content can induce hallucinations in audio-visual LLMs. To systematically study this, we introduce SVHalluc, the first comprehensive benchmark for evaluating speech-vision hallucination in audio-visual LLMs. Our benchmark diagnoses speech-vision hallucinations from two critical and complementary aspects: semantic and temporal. Experimental results demonstrate that state-of-the-art open-source audio-visual LLMs struggle with aligning speech content with corresponding visual signals, with a near-random accuracy on multiple tasks. In contrast, Gemini 2.5 Pro significantly outperforms the open-source models. Our analysis suggests that their failures stem from limited ability in cross-modality understanding, despite strong performance in single-modality perception. Our work uncovers a new and fundamental limitation of current audio-visual LLMs and highlights the need for speech-grounded video comprehension. Project page: https://chenshuang-zhang.github.io/projects/svhalluc/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。