用动态标准自动优化大模型写作,效果媲美人工标注。
Writer-R1: Enhancing Generative Writing in LLMs via Memory-augmented Replay Policy Optimization
- 基于扎根理论构建多智能体协作流程,动态生成可解释的细粒度评价标准。
- 提出MRPO算法,无需额外训练即可实现自我反思与迭代改进。
- 在多个创作任务中超越100B+参数开源模型,适合追求高质量生成的场景。
作为典型的开放式生成任务,创意写作缺乏可验证的参考答案,长期受限于高人力标注成本、评估偏差和粗粒度反馈信号,制约了奖励建模与自动评估。为此,本文首先基于扎根理论设计多智能体协作流程,对问题进行维度分解与层次归纳,动态生成可解释且可复用的细粒度评价标准。进一步提出记忆增强型回放策略优化(MRPO)算法:一方面,在不增加训练的前提下,基于动态标准引导模型开展自我反思,实现可控的迭代优化;另一方面,采用监督微调与强化学习相结合的训练范式,将评价标准转化为奖励信号,实现端到端优化。实验表明,自动生成的标准性能接近人工标注。使用该方法训练的Writer-R1-4B模型在多个创意写作任务中优于基线,并超越部分100B+参数的开源模型。
原文摘要 · Abstract (English)
As a typical open-ended generation task, creative writing lacks verifiable reference answers, which has long constrained reward modeling and automatic evaluation due to high human annotation costs, evaluative bias, and coarse feedback signals. To address these challenges, this paper first designs a multi-agent collaborative workflow based on Grounded Theory, performing dimensional decomposition and hierarchical induction of the problem to dynamically produce interpretable and reusable fine-grained criteria. Furthermore, we propose the Memory-augmented Replay Policy Optimization (MRPO) algorithm: on the one hand, without additional training, MRPO guides models to engage in self-reflection based on dynamic criteria, enabling controlled iterative improvement; on the other hand, we adopt the training paradigm that combines supervised fine-tuning with reinforcement learning to convert evaluation criteria into reward signals, achieving end-to-end optimization. Experimental results demonstrate that the automatically constructed criteria achieve performance gains comparable to human annotations. Writer-R1-4B models trained with this approach outperform baselines across multiple creative writing tasks and surpass some 100B+ parameter open-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。