用可解释的奖励模型提升故事生成创造力,让AI更懂人类审美。
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
- 构建生成式奖励模型,基于教师模型推理链训练并优化。
- 奖励对齐人类判断达68%,显著超越Gemini-2.5-Pro等基线。
- 动态熵激励机制缓解过拟合,适合创意内容生成研究者。
尽管大型语言模型能生成流畅文本,但创作高质量创意故事仍具挑战。强化学习虽具潜力,却面临两大难题:主观故事质量的可靠奖励信号设计,以及训练不稳定性。本文提出面向创意故事生成的强化学习框架(RLCS),系统应对上述问题。首先,我们开发生成式奖励模型(GenRM),通过强教师模型提炼的推理链进行监督微调,并在扩展偏好数据上采用GRPO方法精炼,实现多维度分析与显式推理。其次,引入基于熵的奖励塑形策略,动态优先学习置信度低的错误和不确定的正确预测,防止对已掌握模式的过拟合。实验表明,GenRM在人类创造力判断上的对齐率达68%,且RLCS整体故事质量显著优于包括Gemini-2.5-Pro在内的多个强基线。本工作为强化学习在创意领域应用提供了实用流程,有效克服了奖励建模与训练稳定性的双重挑战。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) can generate fluent text, producing high-quality creative stories remains challenging. Reinforcement Learning (RL) offers a promising solution but faces two critical obstacles: designing reliable reward signals for subjective storytelling quality and mitigating training instability. This paper introduces the Reinforcement Learning for Creative Storytelling (RLCS) framework to systematically address both challenges. First, we develop a Generative Reward Model (GenRM) that provides multi-dimensional analysis and explicit reasoning about story preferences, trained through supervised fine-tuning on demonstrations with reasoning chains distilled from strong teacher models, followed by GRPO-based refinement on expanded preference data. Second, we introduce an entropy-based reward shaping strategy that dynamically prioritizes learning on confident errors and uncertain correct predictions, preventing overfitting on already-mastered patterns. Experiments demonstrate that GenRM achieves 68\% alignment with human creativity judgments, and RLCS significantly outperforms strong baselines including Gemini-2.5-Pro in overall story quality. This work provides a practical pipeline for applying RL to creative domains, effectively navigating the dual challenges of reward modeling and training stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。