噪声训练让神经网络出现麦格克效应,揭示听觉视觉融合机制
Artificial Neural Networks Trained on Noisy Speech Exhibit the McGurk Effect
- 用视听不一致语音测试神经网络,发现其自发产生幻觉音素
- 噪声训练使视觉响应和麦格克效应显著增强,极端噪声则抑制融合
- 无需刻意设计,监督与无监督模型均自然出现该效应,适合认知研究
人类能融合听觉与视觉信息理解语音,这体现在麦格克效应中:当听觉与视觉输入不一致时,会感知到一个虚假的中间音素。本文基于人工神经网络(ANN)研究发展性‘为何’问题的框架,评估了若干在音视频语音上训练的最新模型,测试其对音视频不一致语音的反应。结果显示,即使仅在一致音视频数据上训练的网络,仍表现出麦格克知觉。进一步对比在清晰语音与噪声语音上训练的网络发现,噪声训练显著增强了所有模型的视觉响应和麦格克反应。系统增加训练中的听觉噪声水平可提升视听整合程度,但达到极端噪声水平时,整合能力反而失效。这表明关键期过度噪声暴露可能负面影响视听语音整合的发展。本研究还表明,麦格克效应可在未显式训练的情况下,从监督与无监督网络行为中可靠涌现,支持神经网络作为感知与认知某些方面的有效建模工具。
原文摘要 · Abstract (English)
Humans are able to fuse information from both auditory and visual modalities to help with understanding speech. This is demonstrated through a phenomenon known as the McGurk Effect, during which a listener is presented with incongruent auditory and visual speech that fuse together into the percept of illusory intermediate phonemes. Building on a recent framework that proposes how to address developmental 'why' questions using artificial neural networks, we evaluated a set of recent artificial neural networks trained on audiovisual speech by testing them with audiovisually incongruent words designed to elicit the McGurk effect. We show that networks trained entirely on congruent audiovisual speech nevertheless exhibit the McGurk percept. We further investigated 'why' by comparing networks trained on clean speech to those trained on noisy speech, and discovered that training with noisy speech led to a pronounced increase in both visual responses and McGurk responses across all models. Furthermore, we observed that systematically increasing the level of auditory noise during ANN training also increased the amount of audiovisual integration up to a point, but at extreme noise levels, this integration failed to develop. These results suggest that excessive noise exposure during critical periods of audiovisual learning may negatively influence the development of audiovisual speech integration. This work also demonstrates that the McGurk effect reliably emerges untrained from the behaviour of both supervised and unsupervised networks, even networks trained only on congruent speech. This supports the notion that artificial neural networks might be useful models for certain aspects of perception and cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。