用大模型推理轨迹解析文学质量隐性标准,发现结构比词汇更关键。
What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

- 从大模型推理过程提取文学质量判断逻辑,发现重视作者意图与风格独特性。
- 结构简化导致质量下降2.78分,远超词汇简化仅0.41分的损失。
- 适合研究自动写作评估、计算美学及大模型文本判断机制的人参考。
文学质量的评判标准在文学研究与计算语言学中长期存在争议。本文通过两阶段研究探究具备推理能力的大模型如何评估文学质量。第一阶段构建涵盖六个质量层级的30篇真实文本基准集(从经典文学到匿名论坛帖子),在五次DeepSeek复现中,模型平均分级准确率达79.3%。其推理轨迹揭示出一致的评价理论:更重视意图而非正确性,优先考量文笔、深度与独特文风。对风格匹配但无识别度的仿写文本进行熟悉度实验,发现来源识别可能夸大评分,但此结果与真实质量差异混杂。第二阶段系统破坏五篇经典散文,实施六种操作:词汇简化、节奏扁平化、意象移除、风格泛化、结构简化及综合破坏。结果显示,词汇简化仅造成0.41±0.46分质量损失,显著低于结构简化(2.78)与风格泛化(2.34)。综合破坏造成-5.64分剧降,但呈非叠加效应。与Qwen QwQ的探索性对比显示相似定性模式。整体表明,大模型对文学质量的判断是整体性、作者特定的,且对结构特征更为敏感,对自动写作反馈与计算美学研究具有启示。
原文摘要 · Abstract (English)
What makes writing "good" remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality. In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model's implicit theory of quality from its reasoning traces. Across five DeepSeek replications, the model achieves 79.3% mean tier-classification accuracy. The traces reveal a consistent stated theory: the model values intentionality over correctness, prioritizing craft, depth, and distinctive voice. A familiarity experiment with style-matched but unrecognizable passages suggests that source recognition may inflate scores, although this is confounded by genuine quality differences between canonical originals and researcher-written pastiches. In Study 2, we probe this theory through systematic degradation of five canonical prose passages. We apply six manipulations - vocabulary simplification, rhythm flattening, imagery removal, voice genericization, structure simplification, and combined degradation - and reevaluate each version. Vocabulary simplification causes the smallest quality loss (0.41 +/- 0.46 points), far below structure (2.78) or voice (2.34) loss. Combined degradation is devastating (-5.64) but subadditive. An exploratory comparison with Qwen QwQ shows the same broad qualitative pattern. Together, these studies suggest that LLM judgments of writing quality are holistic, author-specific, and more sensitive to structural than lexical features, with implications for automated writing feedback and computational aesthetics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。