无需参考文本,统一长篇叙事与短篇创意的写作强化学习框架。
UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning
- 自适应约束奖励模型动态生成偏好判断,实现细粒度对齐。
- 无需监督微调,直接优化策略在多任务上显著提升写作质量。
- 模型自发区分需规划与直接生成的任务,具备元认知能力。
创造性写作的核心挑战在于协调长篇叙事的全局连贯性与短篇表达的局部生动性之间的矛盾。长文本生成需要宏观规划,而短文本创作则更依赖即兴、无约束的表达。现有对齐范式通常依赖静态奖励信号和高质量监督数据,成本高且难扩展。为此,我们提出 extbf{UniCreative}:一种统一的无参考强化学习框架。首先引入 extbf{AC-GenRM},一种自适应约束感知的奖励模型,可动态合成与查询相关的评判标准,提供细粒度偏好判断。基于这些信号,提出 extbf{ACPO} 策略优化算法,无需监督微调与真实参考,即可对齐人类偏好,在内容质量与结构范式上均表现优异。实证结果表明,AC-GenRM 与专家评价高度一致,ACPO 在多种写作任务中显著提升性能。关键发现是模型涌现出元认知能力:能自主区分需严谨规划与适合直接生成的任务,验证了直接对齐方法的有效性。
原文摘要 · Abstract (English)
A fundamental challenge in creative writing lies in reconciling the inherent tension between maintaining global coherence in long-form narratives and preserving local expressiveness in short-form texts. While long-context generation necessitates explicit macroscopic planning, short-form creativity often demands spontaneous, constraint-free expression. Existing alignment paradigms, however, typically employ static reward signals and rely heavily on high-quality supervised data, which is costly and difficult to scale. To address this, we propose \textbf{UniCreative}, a unified reference-free reinforcement learning framework. We first introduce \textbf{AC-GenRM}, an adaptive constraint-aware reward model that dynamically synthesizes query-specific criteria to provide fine-grained preference judgments. Leveraging these signals, we propose \textbf{ACPO}, a policy optimization algorithm that aligns models with human preferences across both content quality and structural paradigms without supervised fine-tuning and ground-truth references. Empirical results demonstrate that AC-GenRM aligns closely with expert evaluations, while ACPO significantly enhances performance across diverse writing tasks. Crucially, our analysis reveals an emergent meta-cognitive ability: the model learns to autonomously differentiate between tasks requiring rigorous planning and those favoring direct generation, validating the effectiveness of our direct alignment approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。