arXiv:2605.20364cs.CL2026-05综述

用14维创意评估框架构建了超26万篇长文评述数据集,发现推理监督反而降低生成质量。

When Reasoning Supervision Hurts: TTCW-Based Long-Form Literary Review Generation

论文配图:When Reasoning Supervision Hurts: TTCW-Based Long-Form Literary Review Generation
图 1 · 摘自论文原文
  • 基于TTCW框架构建大规模长文评述数据集,含14维评分与评论
  • 不带推理监督的微调效果更优,最佳得分达0.6820
  • 推理监督易导致重复或无关推理文本,影响固定格式报告完成度

长篇文学写作的自动评估仍具挑战性,通用LLM作为评判者难以捕捉原创性与灵活性等创造性维度。尽管托兰斯写作创造力测试(TTCW)提供了结构化创意评估框架,且已有研究在成对比较中实现基于参考的TTCW评估,但尚无大规模长篇TTCW评述生成数据集。本文填补此空白,构建包含263,911篇长篇故事的数据集,每篇标注14个TTCW维度的标量分值及元合成评注。在此基础上,对Qwen3模型(4B和8B规模)在有无推理内容条件下进行微调。结果表明,无推理微调表现更强且更稳定,最优设置得分达0.6820。进一步分析显示,推理监督模型更易出现解析失败,常持续输出无关或重复的推理文本,而非完成要求的14维度评述报告。这表明,在固定评分量表的评述生成任务中,推理监督并非简单有益,即使经过任务特定微调,精准对齐评分指标仍具挑战。

原文摘要 · Abstract (English)

Automatic evaluation of long-form literary writing remains challenging, as generic LLM-as-Judge approaches may not fully capture creativity-related dimensions such as originality and flexibility. Although the Torrance Test of Creative Writing (TTCW) provides a structured creativity framework, and prior work has demonstrated reference-based TTCW evaluation at the pairwise level, no large-scale dataset exists for long-form TTCW-based literary review generation. We address this gap by constructing a dataset of 263,911 long-form stories, each annotated with scalar scores and meta-synthesised review comments across 14 TTCW-based dimensions. Using this dataset, we fine-tune Qwen3 models at two scales, 4B and 8B, under two conditions: with and without reasoning content. Results show that non-reasoning fine-tuning achieves stronger and more stable performance, with the best setting reaching an evaluation score of 0.6820. Further analysis shows that reasoning-supervised models are more prone to parse failures, often continuing with irrelevant or repetitive reasoning-style text rather than completing the required 14-metric review report. These results suggest that, for fixed-format rubric-based review generation, reasoning supervision is not straightforwardly beneficial, and precise metric-aligned scoring remains challenging even after task-specific fine-tuning.

创意评估长文生成评测基准模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。