arXiv:2601.09609cs.CLcs.AI2026-01ACL被引 4

用多样规划分支提升大模型创作多样性,兼顾质量与创意。

DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing

  • 在思维链中分步规划并主动引入多样性分支
  • 在创作基准上显著提升输出多样性,质量不降
  • 适合需要高创意性的写作任务,如故事生成

基于强化学习(RL)的大语言模型(LLM)在提升生成质量的同时,常导致输出多样性下降,影响其在开放性任务如创意写作中的应用。现有方法缺乏明确的多样性引导机制,更关注优化效率与性能。本文提出一种基于半结构化长思维链(CoT)的强化学习框架,在生成过程中分解为显式规划的中间步骤。引入多样规划分支策略,根据多样性变化在规划阶段主动引入分歧,并设计群体感知多样性奖励,鼓励不同生成轨迹。在创意写作基准上的实验表明,该方法显著提升输出多样性,同时保持生成质量,持续优于现有基线。

原文摘要 · Abstract (English)

Reinforcement learning (RL)-based enhancement of large language models (LLMs) often leads to reduced output diversity, undermining their utility in open-ended tasks like creative writing. Current methods lack explicit mechanisms for guiding diverse exploration and instead prioritize optimization efficiency and performance over diversity. This paper proposes an RL framework structured around a semi-structured long Chain-of-Thought (CoT), in which the generation process is decomposed into explicitly planned intermediate steps. We introduce a Diverse Planning Branching method that strategically introduces divergence at the planning phase based on diversity variation, alongside a group-aware diversity reward to encourage distinct trajectories. Experimental results on creative writing benchmarks demonstrate that our approach significantly improves output diversity without compromising generation quality, consistently outperforming existing baselines.

创意写作强化学习多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。