不同母语者学英语词汇难易度差异,由熟悉度与语言迁移共同决定。
What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty

- 基于母语特征建模词汇难度,融合熟悉度、语义、拼写等多维特征。
- 所有母语者中熟悉度最重要,西班牙语和德语学习者还受拼写迁移影响。
- 模型可解释性强,适合为不同母语者定制词汇教学方案。
我们计算建模了以西班牙语、德语或汉语为母语的学习者在学习英语词汇时的难度。采用梯度提升模型,基于词语的熟悉度(如词频)、语义、表层形式以及跨语言迁移特征进行训练。利用Shapley值分析各特征组的重要性。结果表明,熟悉度是三种母语学习者共有的主导因素。然而,西班牙语和德语学习者的预测还额外依赖于拼写迁移,而汉语学习者因缺乏此类机制,其词汇难度仅由熟悉度与表层特征共同决定。模型提供可解释、针对母语个性化的难度估计,可用于设计针对性词汇课程。
原文摘要 · Abstract (English)
What makes a word difficult to learn, and how does the difficulty depend on the learner's native language? We computationally model vocabulary difficulty for English learners whose first language is Spanish, German, or Chinese with gradient-boosted models trained on features related to a word's familiarity (e.g., frequency), meaning, surface form, and cross-linguistic transfer. Using Shapley values, we determine the importance of each feature group. Word familiarity is the dominant feature group shared by all three languages. However, predictions for Spanish- and German-speaking learners rely additionally on orthographic transfer. This transfer mechanism is unavailable to Chinese learners, whose difficulty is shaped by a combination of familiarity and surface features alone. Our models provide interpretable, L1-tailored difficulty estimates that can be used to design vocabulary curricula.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。