arXiv:2504.03338cs.CL2025-04中稿 · CoNLL 2025被引 6

用词切分探测语言模型的语音表征,跨31种语言验证学习机制。

BabyLM's First Words: Word Segmentation as a Phonological Probing Task

  • 基于预测误差峰值定位词边界,无监督提取模型中的词切分信息。
  • 发现模型即使未在训练中接触词边界,仍能隐式捕捉其位置。
  • 适用于研究儿童语言习得,为子词分词器设计提供实证依据。

语言模型为基于预测的语法理论研究提供了重要框架,但使用大语言模型进行语音分析仍面临挑战:现有语音基准多限于英语,且主流模型输入表示(字符子词)不适于分析音素表征。本文提出将词切分作为语音探针任务,研究在31种语言的儿童导向语料上训练的音素基语言模型所学表征。借鉴词切分计算模型,提出无监督方法,通过观察预测误差在词首处达到峰值来提取词边界。同时使用线性探针验证,这些模型即便未在训练中显式包含词边界,仍能隐式追踪其位置。该跨语言研究支持统计学习理论,并为子词分词器训练方法提供新思路。

原文摘要 · Abstract (English)

Language models provide a key framework for studying linguistic theories based on prediction, but phonological analysis using large language models (LLMs) is difficult; there are few phonological benchmarks beyond English and the standard input representation used in LLMs (subwords of graphemes) is not suitable for analyzing the representation of phonemes. In this work, we demonstrate how word segmentation can be used as a phonological probing task, allowing us to study the representations learned by phoneme-based language models trained on child-directed speech across 31 languages. Following computational models of word segmentation, we present unsupervised methods for extracting word boundaries from a trained model using the observation that prediction-error peaks at the start of words. We also use linear probes to identify that these models implicitly track word boundaries, even when they do not appear in training. This cross-lingual work corroborates statistical learning theories of acquisition and empirically motivates new methods for training subword tokenizers.

语音分析词切分跨语言探针任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。