arXiv:2511.09536cs.CL2025-11被引 2

发现现有可读性度量与自动简化评估和人工判断相关性弱,需明确简化评价的理论框架。

Readability Measures and Automatic Text Simplification: In the Search of a Construct

  • 对比可读性度量、人工判断与自动评估指标之间的相关性
  • 结果显示三者间普遍相关性较低,难以统一评价标准
  • 呼吁建立清晰的简化评价概念体系,适合研究可读性与文本简化者

可读性是海量信息时代的关键概念。为提升文本可读性并使信息对所有人更易获取,自动文本简化(ATS)研究致力于使文本适配目标读者。近期研究关注了自动评估指标与人工判断之间的相关性,但可读性度量(如可读性公式或语言特征)与这两者的关联尚未得到足够重视。本文在英语语境下,通过补充现有评估指标与人工判断的研究,探讨可读性度量在ATS中的作用。首先分析了ATS与可读性研究的关系,随后报告了可读性度量与人工判断、以及可读性度量与ATS评估指标间的相关性研究结果。发现总体上,可读性度量与自动评估指标及人工判断的相关性较弱。我们指出,从不同角度评估简化效果的三类方法之间普遍存在低相关性,因此亟需对ATS中的‘简化’构建进行明确定义。

原文摘要 · Abstract (English)

Readability is a key concept in the current era of abundant written information. To help making texts more readable and make information more accessible to everyone, a line of researched aims at making texts accessible for their target audience: automatic text simplification (ATS). Lately, there have been studies on the correlations between automatic evaluation metrics in ATS and human judgment. However, the correlations between those two aspects and commonly available readability measures (such as readability formulas or linguistic features) have not been the focus of as much attention. In this work, we investigate the place of readability measures in ATS by complementing the existing studies on evaluation metrics and human judgment, on English. We first discuss the relationship between ATS and research in readability, then we report a study on correlations between readability measures and human judgment, and between readability measures and ATS evaluation metrics. We identify that in general, readability measures do not correlate well with automatic metrics and human judgment. We argue that as the three different angles from which simplification can be assessed tend to exhibit rather low correlations with one another, there is a need for a clear definition of the construct in ATS.

可读性文本简化评估指标人机对照

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。