arXiv:2508.06046cs.CLcs.AI2025-08ACL被引 1

用自演化推理提升故事评估,让大模型更懂故事好坏

EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation

  • 通过多人格策略自动生成带评分的思维链数据
  • 在三个评测集上达到当前最佳表现,生成故事质量显著提升
  • 适合需要高质量文本评估与生成的场景

尽管大语言模型作为评判者(LLM-as-a-judge)的有效性已被验证,其在开放式任务中,特别是故事评估方面仍表现有限。准确的故事评估不仅有助于辅助人工判断质量,还能为故事生成提供关键反馈信号。然而,现有方法面临两难:封闭模型依赖提示工程,适应性差;开放模型微调缺乏严谨推理能力。为此,我们提出自演化成对推理框架(EvolvR)。该框架基于成对比较,首先通过多角色策略自合成带评分的思维链(CoT)数据;为保证数据质量,采用多智能体进行自筛选,确保逻辑严密性与鲁棒性;最终以优化后数据训练的评估器作为奖励模型,指导故事生成。实验表明,该框架在StoryER、HANNA和OpenMEVA三个评测集上均达当前最优性能。同时,作为奖励模型使用时,显著提升了生成故事的质量,充分验证了自演化方法的优势。

原文摘要 · Abstract (English)

Although the effectiveness of Large Language Models (LLMs) as judges (LLM-as-a-judge) has been validated, their performance remains limited in open-ended tasks, particularly in story evaluation. Accurate story evaluation is crucial not only for assisting human quality judgment but also for providing key signals to guide story generation. However, existing methods face a dilemma: prompt engineering for closed-source models suffers from poor adaptability, while fine-tuning approaches for open-source models lack the rigorous reasoning capabilities essential for story evaluation. To address this, we propose the Self-Evolving Pairwise Reasoning (EvolvR) framework. Grounded in pairwise comparison, the framework first self-synthesizes score-aligned Chain-of-Thought (CoT) data via a multi-persona strategy. To ensure data quality, these raw CoTs undergo a self-filtering process, utilizing multi-agents to guarantee their logical rigor and robustness. Finally, the evaluator trained on the refined data is deployed as a reward model to guide the story generation task. Experimental results demonstrate that our framework achieves state-of-the-art (SOTA) performance on three evaluation benchmarks including StoryER, HANNA and OpenMEVA. Furthermore, when served as a reward model, it significantly enhances the quality of generated stories, thereby fully validating the superiority of our self-evolving approach.

故事生成模型评估推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。