arXiv:2501.01796cs.CL2025-01

构建易读文本难度数据集,揭示复杂句子的深层原因。

Reading Between the Lines: A dataset and a study on why some texts are tougher than others

  • 基于心理学与翻译研究设计标注体系
  • 用4个预训练模型预测简化策略准确率达78.5%
  • 可解释模型决策,帮助残障人士阅读支持

本研究旨在深入理解为何某些文本对认知功能受限人群(如智商低于70、阅读理解能力弱者)更难阅读。我们提出一种基于心理学实证与翻译研究的难度标注方案,主要利用在线公开的平行文本(标准英语与易读英语版本)构建数据集。通过微调四个预训练Transformer模型,完成多类别分类任务,以预测句子简化所需策略。实验结果显示,最佳模型在测试集上的准确率为78.5%。同时,研究探索了模型决策的可解释性,有助于理解复杂句背后的认知负担。相关资源已开源,地址:https://github.com/Nouran-Khallaf/why-tough。

原文摘要 · Abstract (English)

Our research aims at better understanding what makes a text difficult to read for specific audiences with intellectual disabilities, more specifically, people who have limitations in cognitive functioning, such as reading and understanding skills, an IQ below 70, and challenges in conceptual domains. We introduce a scheme for the annotation of difficulties which is based on empirical research in psychology as well as on research in translation studies. The paper describes the annotated dataset, primarily derived from the parallel texts (standard English and Easy to Read English translations) made available online. we fine-tuned four different pre-trained transformer models to perform the task of multiclass classification to predict the strategies required for simplification. We also investigate the possibility to interpret the decisions of this language model when it is aimed at predicting the difficulty of sentences. The resources are available from https://github.com/Nouran-Khallaf/why-tough

可读性评估无障碍阅读自然语言处理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。