用强化学习让AI写出更具体、可操作的写作建议。
Generating Constructive Feedback on Stories via Reinforcement Learning

- 用多组件奖励函数引导LLM生成针对性反馈。
- 在三个故事数据集上优于Gemini等先进模型。
- 强调可操作建议是反馈有建设性的关键。
建设性反馈对创意写作者提升叙事能力至关重要。由于人工专家反馈成本高、耗时长,大语言模型(LLMs)作为自动写作助手提供了可扩展且高效的替代方案。然而研究表明,当前LLM生成的反馈往往泛化、缺乏可操作性,且难以识别最关键的写作问题。为此,本文提出一种基于强化学习的方法,无需真实反馈即可引导LLM生成建设性反馈。采用群体相对策略优化(GRPO)训练模型,设计新型多组件奖励函数,旨在提升反馈的针对性、有效性及对最核心问题的关注度。在三个故事语料库上的自动与人工评估显示,该方法显著优于现有先进LLM(包括Gemini)和竞争基线。研究发现,提供可操作建议是驱动反馈具有建设性的主要因素。
原文摘要 · Abstract (English)
Constructive feedback is crucial for creative writers to refine their storytelling abilities. Since receiving feedback from human experts is often costly and time-intensive, large language models (LLMs) offer a scalable and efficient alternative as automatic writing assistants. Despite their potential, research indicates that LLM-generated feedback is often generic, lacks actionability, and fails to identify which writing issue is most critical. To address these limitations, we present a reinforcement learning approach that steers LLMs to generate constructive feedback without the need for ground-truth feedback. We train our model using group relative policy optimization (GRPO) with a novel multi-component reward function aiming at constructiveness: it prioritizes feedback that is uniquely tailored to the story, helps to improve story quality, and addresses the most critical writing issue. In automatic and human evaluation across three story corpora, our approach outperforms state-of-the-art LLMs (including Gemini) and competitive baselines. We find that providing actionable suggestions is the main driver of feedback constructiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。