arXiv:2505.11217cs.SDcs.AI2025-05NeurIPS被引 24

AI听声定位常被视觉误导,人类却能靠声音准确定位。

Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization

  • 用3D模拟生成立体音频图像数据,训练新模型EchoPin
  • AI在视听冲突时表现差,人类则依赖声音保持准确
  • 模型呈现类似人类的左右定位偏好,源于立体音频结构

设想听到狗叫转身,却发现只有一辆停着的车,而真正的狗静止在别处。这种感官冲突考验感知能力,但人类能可靠地优先依赖声音。尽管多模态人工智能已融合视觉与听觉,但其如何处理跨模态冲突、是否偏向某一模态仍不清楚。本研究系统考察了AI声音定位中的模态偏差与冲突解决能力,评估主流多模态模型,并在六种音视频条件下(包括一致、冲突、缺失线索)与人类心理物理表现对比。结果显示,人类始终优于AI,对冲突或缺失视觉线索具有更强适应性,依赖听觉信息;而AI模型常默认采用视觉输入,导致性能降至接近随机水平。为此,我们提出受神经科学启发的EchoPin模型,利用3D仿真生成的立体音频-图像数据集进行训练。即便训练数据有限,EchoPin仍超越现有基准。值得注意的是,其表现也呈现出类似人类的水平定位偏倚——更精确区分左右方向,可能源于立体音频结构与人类耳距的对应关系。这些发现表明,感官输入质量与系统架构共同决定多模态表征的准确性。

原文摘要 · Abstract (English)

Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reliably resolve them by prioritizing sound over misleading visuals. Despite advances in multimodal AI integrating vision and audio, little is known about how these systems handle cross-modal conflicts or whether they favor one modality. In this study, we systematically examine modality bias and conflict resolution in AI sound localization. We assess leading multimodal models and benchmark them against human performance in psychophysics experiments across six audiovisual conditions, including congruent, conflicting, and absent cues. Humans consistently outperform AI, demonstrating superior resilience to conflicting or missing visuals by relying on auditory information. In contrast, AI models often default to visual input, degrading performance to near chance levels. To address this, we propose a neuroscience-inspired model, EchoPin, which uses a stereo audio-image dataset generated via 3D simulations. Even with limited training data, EchoPin surpasses existing benchmarks. Notably, it also mirrors human-like horizontal localization bias favoring left-right precision-likely due to the stereo audio structure reflecting human ear placement. These findings underscore how sensory input quality and system architecture shape multimodal representation accuracy.

声音定位多模态认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。