arXiv:2605.07102cs.CL2026-05

用分层框架评估文学质量,让大模型更懂文化、情感与哲思。

SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions

论文配图:SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
图 1 · 摘自论文原文
  • 构建三层解释维度,通过多轮反思和独立验证评估文学作品
  • 98.8%评分一致,跨层相关性达0.65以上,结果高度可靠
  • 可识别生成文本在哲学深度上的短板,适合内容评估研究者

文学质量评估需考察文化表征、情感深度与哲学复杂性等难以量化层面。我们提出SAGE框架,将文学质量分解为基于本体的解释维度,通过结构化大模型评估实现多轮迭代反思与独立验证。在100篇短篇小说(50部经典作品、30篇通俗小说、20篇大模型生成文本)上,按三个分析层(文化、情感心理、存在哲学)进行双模式评估。600次评估中,评分一致性达98.8%,人评一致性超94%,内容与元数据评估模式间近乎完美不变性。统计分析显示显著的体裁层级(经典 > 通俗 > 大模型生成,所有p<0.001),文化批判与哲学深度效应量极大(Cohen's d>2.4),情感表征差距较小(d=1.68),表明情感模式比批判立场或哲思更易从训练数据学习。跨层相关性为r=0.649–0.683,证实三维度可区分质量特征。结果表明,理论驱动的大模型评估具备测量级可靠性,可系统识别生成模型与人类创作的差距,对开放文本生成的自动化评估具有直接意义。

原文摘要 · Abstract (English)

Evaluating literary quality requires assessing interpretive dimensions such as cultural representation, emotional depth, and philosophical sophistication that resist straightforward computational measurement. We introduce SAGE, a hierarchical evaluation framework that decomposes literary quality into ontology-grounded interpretive dimensions assessed through structured large language model evaluation with multi-round iterative reflection and independent validation. We validate the framework on 100 short stories (50 canonical works, 30 pulp fiction, 20 LLM-generated narratives) across three analytical layers (cultural, emotional-psychological, existential-philosophical) using dual-mode assessment. Across 600 evaluations, the framework achieves 98.8% score convergence and greater than 94% inter-rater agreement, with near-perfect mode invariance between content-based and metadata-based evaluation. Statistical analysis reveals a consistent genre hierarchy (Canonical > Pulp > LLM, all p<0.001) with layer-specific discrimination: cultural critique and philosophical depth exhibit very large effect sizes (Cohen's d>2.4), while emotional representation shows smaller gaps (d=1.68), suggesting that affective patterns are more learnable from training data than critical stance or philosophical depth. Cross-layer correlations (r=0.649-0.683) confirm the three dimensions capture empirically distinguishable quality facets. These findings demonstrate that theory-driven LLM evaluation can achieve measurement-grade reliability and support systematic identification of where current generative models fall short of human literary production, with direct implications for scalable automated evaluation of open-ended text generation.

文学评估大模型评测本体论生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。