arXiv:2501.06536cs.CL2025-01

用词频分布范围预测词汇反应时间、熟悉度和复杂度,效果优于传统词频。

Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity

  • 用对数化词频分布范围预测词汇特征,方法简单有效。
  • 在五种语言中,对数范围比词频更优,且能提升词频的预测力。
  • 揭示了语料粒度与对数变换的影响,解释了以往研究矛盾之处。

多种词频分布离散度度量被提出以更全面刻画词语在语料中的分布特征,但外部验证较少。本文在五种不同语言中评估了广泛的离散度度量对词汇反应时间、词汇熟悉度和词汇复杂度的预测能力。结果表明,对数化范围不仅在所有任务和语言中均优于对数词频,而且作为对数词频的补充变量表现最强,持续超越更复杂的离散度度量。文章还探讨了语料部分粒度和对数变换的影响,为以往研究中出现的矛盾结果提供了新解释。

原文摘要 · Abstract (English)

Various measures of dispersion have been proposed to paint a fuller picture of a word's distribution in a corpus, but only little has been done to validate them externally. We evaluate a wide range of dispersion measures as predictors of lexical decision time, word familiarity, and lexical complexity in five diverse languages. We find that the logarithm of range is not only a better predictor than log-frequency across all tasks and languages, but that it is also the most powerful additional variable to log-frequency, consistently outperforming the more complex dispersion measures. We discuss the effects of corpus part granularity and logarithmic transformation, shedding light on contradictory results of previous studies.

词汇认知分布度量语言差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。