arXiv:2601.15869cs.CLcs.AI2026-01

AI语音模型在跨模态感知上逼近人类,但缺乏随机性。

Artificial Rigidities vs. Biological Noise: A Comparative Analysis of Multisensory Integration in AV-HuBERT and Human Observers

  • 用麦格克效应测试模型与人类对视听不一致刺激的反应
  • 模型听觉主导率32.0%接近人类31.8%,但语音融合率达68.0%远超人类47.7%
  • 模型行为严格分类,无人类的感知随机性,适合研究认知差异

本研究通过基准测试(N=44)评估了AV-HuBERT在不一致视听刺激(麦格克效应)下的感知生物保真度。结果揭示出显著的定量同构性:人工智能与人类表现出几乎相同的听觉主导率(32.0% vs. 31.8%),表明模型捕捉到了听觉抵抗的生物阈值。然而,AV-HuBERT显示出对语音融合的确定性偏倚(68.0%),显著高于人类的47.7%。人类表现出感知随机性和多样化的错误模式,而模型始终保持严格分类。研究结果表明,当前自监督架构虽能模拟多感官结果,但缺乏人类语音感知中固有的神经变异性。

原文摘要 · Abstract (English)

This study evaluates AV-HuBERT's perceptual bio-fidelity by benchmarking its response to incongruent audiovisual stimuli (McGurk effect) against human observers (N=44). Results reveal a striking quantitative isomorphism: AI and humans exhibited nearly identical auditory dominance rates (32.0% vs. 31.8%), suggesting the model captures biological thresholds for auditory resistance. However, AV-HuBERT showed a deterministic bias toward phonetic fusion (68.0%), significantly exceeding human rates (47.7%). While humans displayed perceptual stochasticity and diverse error profiles, the model remained strictly categorical. Findings suggest that current self-supervised architectures mimic multisensory outcomes but lack the neural variability inherent to human speech perception.

跨模态感知语音模型神经变异性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。