arXiv:2511.21473cs.CLcs.AI2025-11

用分层网络评估长文档可读性,兼顾句子与全文层级关系。

Hierarchical Ranking Neural Network for Long Document Readability Assessment

  • 双向机制捕捉上下文,定位语义丰富的句子区域。
  • 通过标签减法建模可读性等级顺序,提升预测准确性。
  • 在中英文数据集上表现优于基线模型,适合长文本分析。

可读性评估旨在衡量文本的阅读难度。近年来,尽管深度学习技术逐步应用于该任务,但多数方法未能考虑文本长度或可读性标签的序数关系。本文提出一种双向可读性评估机制,通过捕捉上下文信息识别文本中语义丰富的区域,进而预测句子级可读性标签,并用于辅助推断文档整体可读性水平。此外,引入一对比排序算法,通过标签减法建模可读性等级间的序数关系。在中英文数据集上的实验表明,所提模型性能具有竞争力,优于其他基线模型。

原文摘要 · Abstract (English)

Readability assessment aims to evaluate the reading difficulty of a text. In recent years, while deep learning technology has been gradually applied to readability assessment, most approaches fail to consider either the length of the text or the ordinal relationship of readability labels. This paper proposes a bidirectional readability assessment mechanism that captures contextual information to identify regions with rich semantic information in the text, thereby predicting the readability level of individual sentences. These sentence-level labels are then used to assist in predicting the overall readability level of the document. Additionally, a pairwise sorting algorithm is introduced to model the ordinal relationship between readability levels through label subtraction. Experimental results on Chinese and English datasets demonstrate that the proposed model achieves competitive performance and outperforms other baseline models.

可读性评估长文档分层模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。