揭示语言模型性能与形态学关联研究中的混淆因素
Confounding Factors in Relating Model Performance to Morphology
- 识别实验设计中影响结论的混淆变量
- 发现三种主流假设均受未控变量干扰
- 提出无需人工标注的字节二元组度量方法
语言特征对分词和语言建模的影响程度仍存争议。既有研究认为形态系统影响微乎其微,也有研究强调其关键作用。我们指出这种矛盾源于实验设计中的混淆因素,导致结果难以比较。本文识别出分析形态学与语言建模关系时存在的混淆变量,并重新评估了Arnett & Bergen(2025)提出的三个假设:分词形态对齐性、分词效率和数据集规模。结果表明,这些结论均受未控制变量影响。最后,我们引入字节二元组度量作为内在指标,可无须专家标注预测因果语言建模难度,且与形态复杂度呈梯度相关。研究最终提出了可靠回答形态学与语言建模关系问题的必要条件。
原文摘要 · Abstract (English)
The extent to which individual language characteristics influence tokenization and language modeling is an open question. Differences in morphological systems have been suggested as both unimportant and crucial to consider (Cotterell et al., 2018; Gerz et al., 2018a; Park et al., 2021, inter alia). We argue this conflicting evidence is due to confounding factors in experimental setups, making it hard to compare results and draw conclusions. We identify confounding factors in analyses trying to answer the question of whether, and how, morphology relates to language modeling. Next, we re-assess three hypotheses by Arnett & Bergen (2025) for why modeling agglutinative languages results in higher perplexities than fusional languages: they look at morphological alignment of tokenization, tokenization efficiency, and dataset size. We show that each conclusion includes confounding factors. Finally, we introduce token bigram metrics as an intrinsic way to predict the difficulty of causal language modeling, and find that they are gradient proxies for morphological complexity that do not require expert annotation. Ultimately, we outline necessities to reliably answer whether, and how, morphology relates to language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。