arXiv:2511.09185cs.CL2025-11

用上下文信息衡量文章连贯性,效果优于传统方法。

Context is Enough: Empirical Validation of $\textit{Sequentiality}$ on Essays

  • 仅用上下文词计算连贯性,避免主题选择偏差
  • 在两个数据集上与人工评分相关性更高
  • 结合语言学特征后超越零样本大模型

近期研究提出用大语言模型量化叙事连贯性,采用包含主题和上下文词的顺序性度量。有批评指出原始结果受主题选取干扰,且未与真实连贯性标注验证。本文验证了仅使用上下文词作为更合理、可解释的替代方案。在包含人工标注评分的ASAP++和ELLIPSE两个作文数据集上,上下文版本的顺序性与组织性、连贯性等人评特质得分相关性更强。尽管零样本提示的大模型预测能力更强,但将上下文顺序性与标准语言学特征结合后,预测性能超过零样本大模型,凸显显式建模句间衔接的价值。结果支持使用基于上下文的顺序性作为自动作文评分等任务中可验证、可解释的补充特征。

原文摘要 · Abstract (English)

Recent work has proposed using Large Language Models (LLMs) to quantify narrative flow through a measure called sequentiality, which combines topic and contextual terms. A recent critique argued that the original results were confounded by how topics were selected for the topic-based component, and noted that the metric had not been validated against ground-truth measures of flow. That work proposed using only the contextual term as a more conceptually valid and interpretable alternative. In this paper, we empirically validate that proposal. Using two essay datasets with human-annotated trait scores, ASAP++ and ELLIPSE, we show that the contextual version of sequentiality aligns more closely with human assessments of discourse-level traits such as Organization and Cohesion. While zero-shot prompted LLMs predict trait scores more accurately than the contextual measure alone, the contextual measure adds more predictive value than both the topic-only and original sequentiality formulations when combined with standard linguistic features. Notably, this combination also outperforms the zero-shot LLM predictions, highlighting the value of explicitly modeling sentence-to-sentence flow. Our findings support the use of context-based sequentiality as a validated, interpretable, and complementary feature for automated essay scoring and related NLP tasks.

连贯性评估大模型应用作文评分可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。