arXiv:2604.23860cs.CVcs.AI2026-04中稿 · ICASSP 2026

发现视觉主导的音频幻觉,揭示大模型在第一视角视频中误判声音的问题。

Exploring Audio Hallucination in Egocentric Video Understanding

论文配图:Exploring Audio Hallucination in Egocentric Video Understanding
图 1 · 摘自论文原文
  • 设计问答框架,自动检测第一视角视频中的音频幻觉现象。
  • 测试显示模型对前景与背景音准确率仅27.3%和39.5%。
  • 适合关注多模态模型可靠性与评估方法的研究者。

第一人称视频中,声音是理解用户行为与环境的关键线索,尤其在视觉因持续镜头运动而模糊或遮挡时更为重要。当前先进的视听语言模型(AV-LLMs)虽可生成多模态描述,但本研究揭示其易产生音频幻觉——即从可见但未发声的视觉线索中错误推断出声音。我们构建了一个包含300段第一人称视频的基准数据集,并设计1,000个聚焦声音的问答问题,系统评估模型表现。提出一个基于真实性的分类体系,区分用户动作产生的前景音与环境背景音。评估结果表明,如Qwen2.5 Omni等先进模型在前景音与背景音问答任务上的准确率分别仅为27.3%和39.5%。该工作强调了衡量多模态响应可靠性的重要性,指出鲁棒的幻觉评估对发展可信的AV-LLMs至关重要。

原文摘要 · Abstract (English)

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera movement. State-of-the-art large audio-visual language models (AV-LLMs) can generate multimodal descriptions. However, we show in this work that they are prone to audio hallucinations, often inferring sounds from visual cues that are visible but not heard. We present a systematic and automatic evaluation framework for analyzing audio hallucinations in egocentric video through a targeted question-answering (Q/A) protocol. We curate a dataset of 300 egocentric videos and design 1,000 sound-focused questions to probe model outputs. To characterize hallucinations, we propose a grounded taxonomy that distinguishes between foreground action sounds from the user activities and background ambient sounds. Our evaluation shows that advanced AV-LLMs, such as Qwen2.5 Omni, exhibit high hallucination rates, achieving only 27.3% and 39.5% accuracy on Q/As related to foreground and background sounds, respectively. With this work, we highlight the need to measure the reliability of multimodal responses, emphasizing that robust evaluation of hallucinations is essential to develop reliable AV-LLMs.

音频幻觉多模态第一人称视频模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。