arXiv:2604.19261cs.CL2026-04

用33个语言特征量化评估叙事质量,区分专业与自出版文本

Towards a Linguistic Evaluation of Narratives: A Quantitative Stylistic Framework

  • 提取33个语言特征,分词汇、句法、语义三类进行自动评估
  • 在23本书的语料上,几乎完美区分专业编辑与自出版文本
  • 相比传统指标显著更准,适合自动化故事质量评估场景

叙事质量评估仍具挑战性,因涉及情节、人物塑造和情感影响等主观因素。本文提出一种基于语言维度的量化评估方法,通过提取33个分类为词汇、句法和语义的定量语言特征,实现自动叙事评估。实验在包含23部经典作品与自出版文本的专用语料库上进行。通过相似性矩阵,系统成功聚类叙事,几乎完美地区分了专业编辑与自出版文本。此外,该方法在人类标注数据集上的验证表明,其显著优于传统故事级评估指标,证明了定量语言特征在叙事质量评估中的有效性。

原文摘要 · Abstract (English)

The evaluation of narrative quality remains a complex challenge, as it involves subjective factors such as plot, character development, and emotional impact. This work proposes a quantitative approach to narrative assessment by focusing on the linguistic dimension as a primary indicator of quality. The paper presents a methodology for the automatic evaluation of narrative based on the extraction of a comprehensive set of 33 quantitative linguistic features categorized into lexical, syntactic, and semantic groups. To test the model, an experiment was conducted on a specialized corpus of 23 books, including canonical masterpieces and self-published works. Through a similarity matrix, the system successfully clustered the narratives, distinguishing almost perfectly between professionally edited and self-published texts. Furthermore, the methodology was validated against a human-annotated dataset; it significantly outperforms traditional story-level evaluation metrics, demonstrating the effectiveness of quantitative linguistic features in assessing narrative quality.

叙事评估语言特征自动化评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。