arXiv:2508.19221cs.CL2025-08EMNLP被引 12

传统可读性指标在评估通俗文本时效果差,大模型更接近人类判断。

Evaluating the Evaluators: Are readability metrics good measures of readability?

  • 用大语言模型替代传统可读性公式,更贴近真实阅读难度。
  • 最优模型与人工判断相关性达0.56,传统指标普遍低于0.3。
  • 适合关注通俗化内容质量评估的研究者和应用开发者。

通俗语言摘要(PLS)旨在将复杂文档简化为非专业读者可理解的内容。本文系统调研了PLS领域文献,发现当前可读性评估普遍依赖传统指标如费尔施-金凯德年级水平(FKGL)。然而,这些指标尚未在PLS场景中与人工可读性判断进行对比。我们评估了8种可读性指标,发现多数与人工判断相关性较差,包括最常用的FKGL。进一步实验表明,大语言模型(LMs)作为可读性评价工具表现更优,最佳模型与人工判断的皮尔逊相关系数达0.56。在包含面向非专家的摘要数据集上,LMs更能捕捉深层可读性特征(如所需背景知识),其结论与传统指标存在显著差异。基于此,本文提出可读性评估的最佳实践建议,并开源分析代码与调查数据。

原文摘要 · Abstract (English)

Plain Language Summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences. In this paper, we conduct a thorough survey of PLS literature, and identify that the current standard practice for readability evaluation is to use traditional readability metrics, such as Flesch-Kincaid Grade Level (FKGL). However, despite proven utility in other fields, these metrics have not been compared to human readability judgments in PLS. We evaluate 8 readability metrics and show that most correlate poorly with human judgments, including the most popular metric, FKGL. We then show that Language Models (LMs) are better judges of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments. Extending our analysis to PLS datasets, which contain summaries aimed at non-expert audiences, we find that LMs better capture deeper measures of readability, such as required background knowledge, and lead to different conclusions than the traditional metrics. Based on these findings, we offer recommendations for best practices in the evaluation of plain language summaries. We release our analysis code and survey data.

可读性评估大模型应用通俗语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。