arXiv:2410.02691cs.CL2024-10EMNLP被引 18

将分词模型转化为字符级模型,提升语言认知研究的准确性

On the Proper Treatment of Tokenization in Psycholinguistics

  • 用字符级语言模型替代分词级模型,解决区域与分词对不齐问题
  • 实证发现多个新焦点区域的意外度预测效果优于原研究区域
  • 适用于语言认知、阅读行为等心理学实验设计者

语言模型广泛用于计算心理语言学中,通过语言模型对目标区域(字符子串)的负对数概率(即意外度)来推断读者的认知负荷。然而,现代语言模型在训练中需经过分词步骤,导致模型实际建模的是分词序列而非字符序列,而研究中的目标区域通常与分词边界不一致,造成严重偏差。本文提出:在使用语言模型进行心理语言学分析前,应将分词级语言模型近似地边缘化为字符级语言模型,从而能准确计算任意字符子串(称为焦点区域)的意外度。该方法独立于具体分词方案,有效解决了对齐问题。实证表明,多个新的焦点区域其意外度比原始目标区域的预测能力更强。

原文摘要 · Abstract (English)

Language models are widely used in computational psycholinguistics to test theories that relate the negative log probability (the surprisal) of a region of interest (a substring of characters) under a language model to its cognitive cost experienced by readers, as operationalized, for example, by gaze duration on the region. However, the application of modern language models to psycholinguistic studies is complicated by the practice of using tokenization as an intermediate step in training a model. Doing so results in a language model over token strings rather than one over character strings. Vexingly, regions of interest are generally misaligned with these token strings. The paper argues that token-level language models should be (approximately) marginalized into character-level language models before they are used in psycholinguistic studies to compute the surprisal of a region of interest; then, the marginalized character-level language model can be used to compute the surprisal of an arbitrary character substring, which we term a focal area, that the experimenter may wish to use as a predictor. Our proposal of marginalizing a token-level model into a character-level one solves this misalignment issue independently of the tokenization scheme. Empirically, we discover various focal areas whose surprisal is a better psychometric predictor than the surprisal of the region of interest itself.

语言模型认知科学意外度分词对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。