arXiv:2604.24470cs.CL2026-04ACL被引 2

用大模型零样本评估文本可读性,效果优于传统方法。

Zero-shot Large Language Models for Automatic Readability Assessment

  • 设计新提示策略,无需训练即可评估可读性。
  • 在14个数据集上,新方法13个表现更优。
  • 融合上下文与句长等特征,跨语言更稳定。

无监督自动可读性评估(ARA)在医疗和教育内容适配中具有重要应用价值。本文提出一种新的零样本提示方法,并首次系统评估10种不同规模与来源的开源大语言模型(LLMs)在14个多样化数据集上的表现(涵盖不同文本长度与语言)。结果表明,所提方法在13个数据集中超越已有方法。此外,我们提出LAURAE,结合大模型与可读性公式评分,同时捕捉上下文信息与句长等浅层特征,显著提升鲁棒性。评估显示,LAURAE在多种语言、文本长度及技术术语密度下均优于现有方法。

原文摘要 · Abstract (English)

Unsupervised automatic readability assessment (ARA) methods have important practical and research applications (e.g., ensuring medical or educational materials are suitable for their target audiences). In this paper, we propose a new zero-shot prompting methodology for ARA and present the first comprehensive evaluation of using large language models (LLMs) as an unsupervised ARA method by testing 10 diverse open-source LLMs (e.g., different sizes and developers) on 14 diverse datasets (e.g., different text lengths and languages). Our findings show that our proposed prompting methodology outperforms prior methods on 13 of the 14 datasets. Furthermore, we propose LAURAE, which combines LLM and readability formula scores to improve robustness by capturing both contextual and shallow (e.g., sentence length) features of readability. Our evaluation demonstrates that LAURAE robustly outperforms prior methods across languages, text lengths, and amounts of technical language.

可读性评估大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。