arXiv:2509.22973cs.CL2025-09EMNLP被引 4

自监督语音模型通过几何结构隐式学习词形变化规律,无需显式语言单元。

Emergent morpho-phonological representations in self-supervised speech models

  • 利用全局线性几何结构建模词形变化关系
  • 能准确映射英语名词动词的规则变形形式
  • 适合研究语音认知机制与模型可解释性

自监督语音模型可在自然、嘈杂环境中高效识别口语词汇,但我们尚不清楚它们使用何种语言表征完成此任务。为此,我们研究了针对词识别优化的S3M变体在英语常见名词和动词屈折中的音位与形态表征。发现其表征具有全局线性几何结构,可用于连接英语名词与动词与其规则屈折形式。该几何结构并不直接追踪音位或形态单位,而是捕捉英语词库中大量词对间的分布关系——通常但非总是源于形态屈折。这些发现揭示了可能支持人类口语词识别的候选表征策略,挑战了音位与形态需独立表征的假设。

原文摘要 · Abstract (English)

Self-supervised speech models can be trained to efficiently recognize spoken words in naturalistic, noisy environments. However, we do not understand the types of linguistic representations these models use to accomplish this task. To address this question, we study how S3M variants optimized for word recognition represent phonological and morphological phenomena in frequent English noun and verb inflections. We find that their representations exhibit a global linear geometry which can be used to link English nouns and verbs to their regular inflected forms. This geometric structure does not directly track phonological or morphological units. Instead, it tracks the regular distributional relationships linking many word pairs in the English lexicon -- often, but not always, due to morphological inflection. These findings point to candidate representational strategies that may support human spoken word recognition, challenging the presumed necessity of distinct linguistic representations of phonology and morphology.

语音模型表征学习形态学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。