用GPT-4评估并调节英文文本可读性,效果接近人工判断。
Measuring and Modifying the Readability of English Texts with GPT-4
- 零样本调用GPT-4 Turbo和GPT-4o mini评估可读性
- 与人工判断相关性达0.74~0.76,优于传统公式
- 可有效让文本变易或变难读,但人类判断仍有差异
大型语言模型在其他领域取得成功,引发了其能否可靠评估与调控文本可读性的疑问。我们通过实证方法研究该问题。首先,基于包含4,724个英文文本片段的公开语料库,发现从GPT-4 Turbo和GPT-4o mini零样本生成的可读性评分与人工判断高度相关(相关系数分别为r = 0.76和r = 0.74),优于传统可读性公式及多种心理语言学指标。随后,在一项预注册的人类实验中(N = 59),我们检验GPT-4 Turbo是否能稳定地使文本更易或更难阅读。结果支持该假设,但人类判断中的部分方差仍未被解释。最后讨论了该方法的局限性,包括适用范围有限,以及‘可读性’概念本身对上下文、受众和目标的依赖性。
原文摘要 · Abstract (English)
The success of Large Language Models (LLMs) in other domains has raised the question of whether LLMs can reliably assess and manipulate the readability of text. We approach this question empirically. First, using a published corpus of 4,724 English text excerpts, we find that readability estimates produced ``zero-shot'' from GPT-4 Turbo and GPT-4o mini exhibit relatively high correlation with human judgments (r = 0.76 and r = 0.74, respectively), out-performing estimates derived from traditional readability formulas and various psycholinguistic indices. Then, in a pre-registered human experiment (N = 59), we ask whether Turbo can reliably make text easier or harder to read. We find evidence to support this hypothesis, though considerable variance in human judgments remains unexplained. We conclude by discussing the limitations of this approach, including limited scope, as well as the validity of the ``readability'' construct and its dependence on context, audience, and goal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。