发现视觉模型常误判声音,提出新方法验证音频真实性的检测框架。
When Vision Speaks for Sound

- 设计三种反事实音频干预:错位、静音、互换,测试模型是否真听懂音频。
- 通过干预训练,模型在三类测试上性能提升28个百分点,且不损害通用能力。
- 适合研究多模态对齐、音频视觉融合的学者,尤其关注模型可信度者。
尽管视频多模态大模型发展迅速,我们发现其看似具备音频理解能力,实则主要依赖视觉线索推断或虚构声学信息,而非验证音频流本身。该问题在开源与闭源领先模型中普遍存在。我们将其归因于音频-视觉‘聪明汉斯效应’——模型看似音频对齐,实则利用视觉与声音的关联性,却不验证二者是否真正同步。为此,我们提出Thud框架,基于三种反事实音频编辑:时移(Shift)测试时间同步性,静音(Mute)测试声音存在性,交换(Swap)测试音画一致性。在诊断基础上,进一步提出两阶段对齐方案:以干预生成的偏好对训练音频验证能力,事件级通用视频偏好防止过拟合。最优10,000样本方案使三项干预测试平均性能提升28个百分点,同时轻微提升通用视频与音画问答基准表现。
原文摘要 · Abstract (English)
Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate acoustic information, rather than verifying the audio stream. This issue appears across both state-of-the-art open-source omni models and leading closed-source models from providers such as Google and OpenAI. We characterize this failure mode as an audio-visual Clever Hans effect, in which models appear (falsely) audio-grounded, but actually exploit visual-acoustic correlations without verifying whether the audio and visual streams are truly aligned. To systematically study this behavior, we introduce Thud, an intervention-driven probing framework based on three counterfactual audio edits: Shift, which tests temporal synchronization; Mute, which tests sound existence; and Swap, which tests audio-visual consistency. Beyond diagnosis, we further study a two-stage alignment recipe: intervention-derived preference pairs teach audio verification, while event-level general video preferences regularize the model against over-specialization. Our best 10K-sample recipe improves average performance across the three intervention dimensions by 28 percentage points, while slightly improving performance on general video and audio-visual QA benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。