AI语音模型在跨模态感知上逼近人类,但缺乏随机性。
Artificial Rigidities vs. Biological Noise: A Comparative Analysis of Multisensory Integration in AV-HuBERT and Human Observers
- 用麦格克效应测试模型与人类对视听不一致刺激的反应
- 模型听觉主导率32.0%接近人类31.8%,但语音融合率达68.0%远超人类47.7%
- 模型行为严格分类,无人类的感知随机性,适合研究认知差异
本研究通过基准测试(N=44)评估了AV-HuBERT在不一致视听刺激(麦格克效应)下的感知生物保真度。结果揭示出显著的定量同构性:人工智能与人类表现出几乎相同的听觉主导率(32.0% vs. 31.8%),表明模型捕捉到了听觉抵抗的生物阈值。然而,AV-HuBERT显示出对语音融合的确定性偏倚(68.0%),显著高于人类的47.7%。人类表现出感知随机性和多样化的错误模式,而模型始终保持严格分类。研究结果表明,当前自监督架构虽能模拟多感官结果,但缺乏人类语音感知中固有的神经变异性。
原文摘要 · Abstract (English)
This study evaluates AV-HuBERT's perceptual bio-fidelity by benchmarking its response to incongruent audiovisual stimuli (McGurk effect) against human observers (N=44). Results reveal a striking quantitative isomorphism: AI and humans exhibited nearly identical auditory dominance rates (32.0% vs. 31.8%), suggesting the model captures biological thresholds for auditory resistance. However, AV-HuBERT showed a deterministic bias toward phonetic fusion (68.0%), significantly exceeding human rates (47.7%). While humans displayed perceptual stochasticity and diverse error profiles, the model remained strictly categorical. Findings suggest that current self-supervised architectures mimic multisensory outcomes but lack the neural variability inherent to human speech perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。