arXiv:2502.11150cs.CL2025-02被引 1

用眼动追踪评估文本易读性,发现现有方法预测能力差。

Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods

  • 通过眼动数据捕捉实时阅读难度,构建认知评估框架。
  • 主流方法对成人阅读难易度的预测远不如心理语言学中的词属性。
  • 适用于教育、NLP等领域研究者,推动更贴近真实阅读体验的评估方法。

自动文本可读性评分方法研究已有百年历史,广泛应用于科研与各类用户场景。以往主要依赖阅读理解测试成绩和可读性评分两类离线人类行为数据。本文转而关注可读性中基础但被忽视的实时阅读难易度,采用眼动追踪技术获取在线阅读数据,提出一种新的认知评估框架,量化不同评分方法对阅读难易度的解释能力,同时控制文本内容差异。将该框架应用于传统可读性公式、基于NLP的方法、教育领域商用系统及前沿大模型,结果表明:这些方法在预测英语成人阅读难易度方面均表现不佳,远逊于心理语言学中常用于预测阅读时间的词属性。这一结论在母语与第二语言读者、不同阅读方式及不同文本长度下均成立。研究揭示了广泛可读性方法的重要局限,凸显实时行为基准的价值,并呼吁发展更符合认知规律的可读性评分新范式。

原文摘要 · Abstract (English)

Automatic methods for scoring text readability have been studied for over a century, and are widely used in research and in user-facing applications in many domains. Thus far, the development and evaluation of such methods have primarily relied on two types of offline human behavioral data, performance on reading comprehension tests and ratings of text readability levels. In this work, we instead focus on a fundamental and understudied aspect of readability, real-time reading ease, captured with online reading measures using eye tracking. We introduce a new cognitive evaluation framework for readability scoring methods that quantifies their ability to account for reading ease, while controlling for content variation across texts. Applying this evaluation to prominent traditional readability formulas, NLP-based methods, commercial systems used in education, and frontier LLMs suggests that they are all poor predictors of English reading ease in adults as compared to word properties commonly used in psycholinguistics for the prediction of reading times. This outcome holds across L1 and L2 speakers, different reading regimes, and textual units of different lengths. Our results reveal an important limitation of a wide range of methods for readability scoring, highlight the utility of real-time behavioral benchmarks for readability research, and call for new, cognitively driven readability scoring approaches that can better account for how humans experience texts in real time.

可读性评估眼动追踪认知计算NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。