提出六层框架评估叙事的文学解读质量,突破现有指标局限。
SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives

- 分层设计:规则判断文本表征+大模型评估文化、情感、哲思等深层维度
- 多轮迭代验证实现98.8%收敛与超94%一致性,跨模型稳定可靠
- 发现情绪心理表达近人类水平,文化批判与哲学深度仍有双倍差距
评估叙事的文学质量需衡量文化表征、情感深度和哲学参与等诠释性维度,现有NLG评估指标难以覆盖。我们提出SAGE,一种六层评估框架,将可观测文本属性的规则化评估与基于文化理论、情感理论和存在主义哲学的大型语言模型(LLM)评估相分离。每个诠释层均通过多轮迭代式LLM评估与独立交叉验证,实现测量级可靠性(98.8%收敛率,>94%评分者间一致率),且在不同评估模型间保持稳定。在100篇短故事共600次评估中,核心发现为系统性能力边界:情绪-心理表征接近人类水平,而文化批判与哲学深度差距约为两倍。生成叙事在所有三层均低于商业类型小说表现。我们将其解释为:模式可复制的文学能力可从训练语料中学得,但立场性要求的文化定位与哲学参与,仅靠模式匹配无法实现。
原文摘要 · Abstract (English)
Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representation, emotional depth, and philosophical engagement) that existing NLG metrics cannot measure. We introduce SAGE, a six-layer evaluation framework that separates rule-based assessment of observable textual properties from LLM-based evaluation of interpretive qualities drawn from cultural theory, affect theory, and existentialist philosophy. Each interpretive layer is assessed through multi-round iterative LLM evaluation with independent cross-validation, achieving measurement-grade reliability (98.8% convergence, >94% inter-rater agreement) stable across evaluator models. Validated on 600 evaluations across 100 short stories, our central finding is a systematic capability boundary: emotional-psychological representation approaches human levels, while cultural critique and philosophical depth exhibit approximately double the gap. LLM-generated narratives score below even commercial genre fiction on all three layers. We interpret this as a boundary between pattern-reproducible literary capacities learnable from training corpora and stance-requiring ones demanding cultural positioning and philosophical engagement that pattern matching alone cannot provide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。