发现文本信息量和主题比字面特征更影响可读性,模型类指标表现显著优于传统方法。
Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
- 通过897份人类判断分析可读性影响因素,发现信息量与主题是关键
- 15个传统指标平均排名8.6,4个模型基指标始终位居前四
- 适合关注可读性评估改进的研究者或内容创作者
自动可读性评估对确保有效且易懂的书面交流至关重要。尽管已有进展,但该领域仍受可读性定义不一及依赖表面文本特征的测量方式制约。本文通过对897份人类可读性判断的跨数据集分析,发现除表层线索外,信息含量与主题显著影响文本可理解性。我们进一步在五个英文数据集上评估了15种主流可读性指标,并与六种更精细的模型基指标对比。结果表明,四种模型基指标在与人类判断的等级相关性中始终位列前四,而表现最佳的传统指标平均排名为8.6。研究揭示当前可读性指标与人类感知存在明显偏差,凸显模型基方法更具前景。
原文摘要 · Abstract (English)
Automatic readability assessment plays a key role in ensuring effective and accessible written communication. Despite significant progress, the field is hindered by inconsistent definitions of readability and measurements that rely on surface-level text properties. In this work, we investigate the factors shaping human perceptions of readability through the analysis of 897 judgments, finding that, beyond surface-level cues, information content and topic strongly shape text comprehensibility. Furthermore, we evaluate 15 popular readability metrics across five English datasets, contrasting them with six more nuanced, model-based metrics. Our results show that four model-based metrics consistently place among the top four in rank correlations with human judgments, while the best performing traditional metric achieves an average rank of 8.6. These findings highlight a mismatch between current readability metrics and human perceptions, pointing to model-based approaches as a more promising direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。