TF-IDF的本质可由词频突现的假设检验推导,揭示其统计学原理。
Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness
- 用带伽马惩罚的贝塔-二项分布建模词频突现,对比二项分布的零假设。
- 推导出的词权重方案在文本分类上表现接近TF-IDF。
- 为词权重设计提供新的统计检验框架,适合自然语言处理研究者。
TF-IDF是一种广泛用于识别文档中重要词汇的经典公式。本文表明,类似TF-IDF的得分可自然地从一个捕捉词频突现(即词频过度离散)的惩罚似然比检验的检验统计量中导出。在该框架中,备择假设通过一组具有伽马惩罚项的贝塔-二项分布来建模文档集合中的词频分布,以捕捉词频突现;而零假设则假设词频服从二项分布,无法反映词频突现现象。我们发现,由该检验统计量导出的词权重方案在文档分类任务上的表现与TF-IDF相当。本文从统计学角度深化了对TF-IDF的理解,并强调了假设检验框架在推进词权重设计方面的潜力。
原文摘要 · Abstract (English)
TF-IDF is a classical formula that is widely used for identifying important terms within documents. We show that TF-IDF-like scores arise naturally from the test statistic of a penalized likelihood-ratio test setup capturing word burstiness (also known as word over-dispersion). In our framework, the alternative hypothesis captures word burstiness by modeling a collection of documents according to a family of beta-binomial distributions with a gamma penalty term on the precision parameter. In contrast, the null hypothesis assumes that words are binomially distributed in collection documents, a modeling approach that fails to account for word burstiness. We find that a term-weighting scheme given rise to by this test statistic performs comparably to TF-IDF on document classification tasks. This paper provides insights into TF-IDF from a statistical perspective and underscores the potential of hypothesis testing frameworks for advancing term-weighting scheme development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。