调整词元粒度能显著影响语言模型困惑度对阅读难度的预测能力。
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
- 通过控制词元粒度,研究其对语言模型困惑度的影响。
- 8000个词元词汇量时,困惑度对自然文本阅读时间的预测力最强。
- 粗粒度词元更敏感地捕捉到歧义句式的处理困难,适合认知建模。
单词级语言模型困惑度常被用于模拟人类读者的增量加工过程,但语言建模中诸多选择如何影响其预测能力仍存疑问。其中一个被忽视的因素是子词词元的粒度,它显式编码了词长和频率信息,最终影响所学向量表示的质量。本文通过操纵词元粒度并评估其对自然文本及歧义句式处理难度预测能力的影响。实验显示,在自然阅读时间任务中,8000个词元词汇量下的困惑度预测力最强;而在歧义句式任务中,使用粗粒度词元的语言模型通常在关键位置赋予更高困惑度,表现出比以往研究更强的歧义敏感性。结果表明,词元粒度在语言模型困惑度的认知建模质量中起着重要作用。
原文摘要 · Abstract (English)
Word-by-word language model surprisal is often used to model the incremental processing of human readers, which raises questions about how various choices in language modeling influence its predictive power. One factor that has been overlooked in cognitive modeling is the granularity of subword tokens, which explicitly encodes information about word length and frequency, and ultimately influences the quality of vector representations that are learned. This paper presents experiments that manipulate the token granularity and evaluate its impact on the ability of surprisal to account for processing difficulty of naturalistic text and garden-path constructions. Experiments with naturalistic reading times reveal a substantial influence of token granularity on surprisal, with tokens defined by a vocabulary size of 8,000 resulting in surprisal that is most predictive. In contrast, on garden-path constructions, language models trained on coarser-grained tokens generally assigned higher surprisal to critical regions, suggesting a greater sensitivity to garden-path effects than previously reported. Taken together, these results suggest a large role of token granularity on the quality of language model surprisal for cognitive modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。