长文本可读性得分主要由主题分布决定,与词汇细节无关。
Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models
- 在主题模型下,可读性分数仅依赖文档主题分布,不受词汇组成影响。
- 实验显示,一半文本的主题能预测另一半的可读性得分,相关性达0.779至0.884。
- 结果适用于文本分析、主题建模研究者,但不解释人类阅读机制。
Flesch Reading Ease (FRE) 和 Flesch-Kincaid Grade Level (FKGL) 是基于相同两个文档统计量计算的英语可读性评分,但其在长文档中的稳定性并不意味着对词汇组成的不变性。令人惊讶的是,在具有显式句边界标记的主题模型下,这两个分数在长文本极限中几乎必然收敛为文档主题分布的确定性函数,仅通过两个标量率中介。所有分数变化均由主题构成决定,而非残留的可读性信号。该理论涵盖两种公式,实验评估了 FKGL。在固定混合比例且秩为 [1, q, s] = 3 的情况下,通过内部主题向量的纤维是局部 (K-3) 维的,而等分值水平集是局部 (K-2) 维且弯曲的。在布朗语料库和书面BNC两个平衡语料库的交叉验证中,从文档一半内容词推断出的主题向量,对另一半的 FKGL 预测相关系数分别为 0.779 和 0.884。在布朗语料库上,加入主题预测后,模型决定系数增量 ΔR² = 0.002,置信区间包含零;在 BNC 上,对应增量为 0.024,五次 K=100 拟合中有四次为正(中位数 0.021)。由于推断主题可能吸收体裁、语域和风格因素,这些结果不用于解释人类可读性或因果效应。
原文摘要 · Abstract (English)
Flesch Reading Ease (FRE) and the Flesch-Kincaid Grade Level (FKGL) are widely used readability scores for English computed from the same two document statistics, yet their stability on long documents need not imply invariance to lexical composition. Surprisingly, under a topic model with an explicit sentence-boundary token, both scores converge almost surely to deterministic functions of the document topic distribution through just two scalar rates: in the long-text limit, all score variation is mediated by topical composition rather than any residual readability signal. The theory covers both formulae, while the experiments evaluate FKGL. In a fixed admixture with rank[1, q, s] = 3, fibres through interior topic vectors are locally (K-3)-dimensional, whereas regular iso-score level sets are locally (K-2)-dimensional and curved. In out-of-fold evaluation on two balanced corpora, Brown and the written BNC, a topic vector inferred from one document half's content words predicts the other half's FKGL at r = 0.779 and 0.884, respectively. On Brown, adding the topic prediction to genre and mean content-word syllable count yields $ΔR^2$ = 0.002, with a confidence interval spanning zero; on the BNC, the corresponding split-half increment is 0.024, positive in four of five K = 100 fits (median 0.021). Because inferred topics may also absorb genre, register, and style, we do not interpret these results as evidence about human readability or causal effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。