让AI写作像人一样反思修改,显著提升创意文本质量。
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
- 通过迭代写作者-评审者交互,生成带反思与修订的高质量思考路径。
- 在多个创作与深度研究任务上,性能显著优于传统模型。
- 适合需要高质量创意内容生成的研究者与内容创作者。
尽管长链推理在数学等可验证领域大幅提升大模型表现,但在开放性写作任务中的效果仍不明确。本文系统研究发现,现有主流推理模型在开放性写作任务中提升有限,原因在于缺乏深层反思与修订机制,导致改进幅度远低于数学推理任务。为此,我们提出R2-Write:一种自动化框架,通过迭代写作者-评审者交互,生成富含显式反思与修订模式的高质量思维轨迹。为避免冗余反思,设计过程奖励机制,在强化学习中监督反思质量,同时提升性能与生成效率。在多个创意写作与深度研究基准上的实验表明,显式引入反思与修订模式能有效激活开放性写作任务中的深度推理能力。
原文摘要 · Abstract (English)
While deep reasoning with long chain-of-thought has dramatically improved large language models in verifiable domains like mathematics, its effectiveness for open-ended tasks such as writing remains unexplored. In this paper, we conduct a systematic investigation revealing that existing mainstream reasoning models achieve limited gains on open-ended writing tasks. Our further analysis shows that these models lack deep reflection and revision patterns in open-ended writing, resulting in substantially smaller improvements compared to mathematical reasoning tasks. To address this limitation, we introduce R2-Write: an automated framework that synthesizes high-quality thinking trajectories enriched with explicit reflection and revision patterns through iterative writer-judge interaction. To prevent redundant reflections, we design a process reward mechanism that supervises reflection quality during reinforcement learning, improving both performance and token efficiency. Extensive experiments across multiple creative writing and deep-research benchmarks demonstrate significant improvements, validating that explicitly incorporating reflection and revision patterns unlocks deep reasoning capabilities for open-ended writing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。