arXiv:2410.18590cond-mat.stat-mechcs.CL2024-10

用动态吸引子模拟语音识别,解释了为何短词易辨识、长词易混淆。

Speech perception: a model of word recognition

  • 用下降动力学建模语音,单词对应吸引子状态。
  • 短词识别快且准确,长词有陷入歧义的可能。
  • 适合研究语音感知机制或语言认知的心理学家。

我们提出一种语音感知模型,考虑了音素间的相关性。在该模型中,单词对应于适当选择的下降动力学的吸引子。由此产生的词汇表中短词丰富,长词较少,符合合理的词长分布。我们分别考察了在误听情况下的短词与长词解码过程。在短词情况下,算法要么快速恢复正确单词,要么提出另一个有效单词;而在长词情况下,虽然成功解码仍相对较快,但存在有限概率永久迷失于合适单词的景观中,无法稳定收敛。

原文摘要 · Abstract (English)

We present a model of speech perception which takes into account effects of correlations between sounds. Words in this model correspond to the attractors of a suitably chosen descent dynamics. The resulting lexicon is rich in short words, and much less so in longer ones, as befits a reasonable word length distribution. We separately examine the decryption of short and long words in the presence of mishearings. In the regime of short words, the algorithm either quickly retrieves a word, or proposes another valid word. In the regime of longer words, the behaviour is markedly different. While the successful decryption of words continues to be relatively fast, there is a finite probability of getting lost permanently, as the algorithm wanders round the landscape of suitable words without ever settling on one.

语音识别认知模型动态系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。