arXiv:2507.00769cs.CLcs.AI2025-07Conference of the …被引 39

首个可信赖的创意写作评估基准,解决AI写故事难量化问题。

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

  • 构建2480组人类标注的叙事对比数据集,用于验证AI写作质量。
  • 训练的奖励模型准确率达78%,优于所有现成大模型裁判。
  • 适合研究生成式内容评价、人机偏好对齐的学者与开发者。

评估大型语言模型生成的创意写作仍具挑战性,因开放性叙事缺乏真实答案。当前常使用现成语言模型作为零样本裁判,但其可靠性未明。为此,我们推出LitBench,首个标准化创意写作评估基准与配套数据集,包含从Reddit获取的2,480组去偏的人类标注故事对比测试集,以及43,827对人类偏好标签的训练语料。基于LitBench,我们(一)评测零样本大模型裁判,(二)训练布拉德利-特里与生成式奖励模型,(三)开展在线人类实验验证奖励模型在新生成故事中的排名一致性。结果表明,Claude-3.7-Sonnet为最强现成裁判,与人类偏好达成73%一致;训练后的奖励模型均达78%准确率,优于所有现成裁判。在线人类实验进一步确认其在新生成故事中持续符合人类偏好。我们已在Hugging Face发布LitBench与奖励模型,为可靠自动化评估与优化创意写作系统提供经验证资源。

原文摘要 · Abstract (English)

Evaluating creative writing generated by large language models (LLMs) remains challenging because open-ended narratives lack ground truths. Without performant automated evaluation methods, off-the-shelf (OTS) language models are employed as zero-shot judges, yet their reliability is unclear in this context. In pursuit of robust evaluation for creative writing, we introduce LitBench, the first standardized benchmark and paired dataset for creative writing verification, comprising a held-out test set of 2,480 debiased, human-labeled story comparisons drawn from Reddit and a 43,827-pair training corpus of human preference labels. Using LitBench, we (i) benchmark zero-shot LLM judges, (ii) train Bradley Terry and generative reward models, and (iii) conduct an online human study to validate reward model rankings on newly LLM-generated stories. Our benchmark identifies Claude-3.7-Sonnet as the strongest off-the-shelf judge, reaching 73% agreement with human preferences; among trained reward models, Bradley-Terry and Generative reward models both attain an accuracy of 78%, outperforming all off-the-shelf judges. An online human study further confirms that our trained reward models consistently align with human preferences in novel LLM-generated stories. We release LitBench and reward models at https://huggingface.co/collections/SAA-Lab/litbench-68267b5da3aafe58f9e43461, providing a vetted resource for reliable, automated evaluation and optimization of creative writing systems.

创意写作评估基准奖励模型人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。