arXiv:2605.27832cs.CL2026-05

用词类游戏训练大模型,让创造力可衡量、可提升。

Playing with Words, Improving with Rewards: Training Language Models for Creative Association

  • 用密语游戏训练模型,同时锻炼发散与收敛思维。
  • 8B模型在10项创造力任务中8项提升,推理仅轻微下降。
  • 适合想提升模型创意能力的研究者和开发者。

大型语言模型正被应用于越来越复杂的任务,需要具备创造性以有效探索庞大的解空间。然而,创造力的主观性及人类判断的局限性使得训练具有创造力的模型尤为困难。为此,我们基于密语游戏(Codenames)训练语言模型,该游戏能同时锻炼创造性的两个核心维度——发散思维与收敛思维,并提供客观可验证的结果。这种可验证性使我们能够绕过人工评判,采用可验证奖励的强化学习(RLVR)进行训练。我们训练了Qwen3-1.7B、4B和8B三个规模的模型,并在10个创造力基准和4个推理基准上进行评估。结果表明,精度与多样性之间的权衡具有规模依赖性:8B模型更注重创造力,而1.7B和4B模型则在牺牲部分创造力的前提下显著提升了推理精度。具体而言,8B模型在10项创造力任务中有8项表现提升,推理能力仅有轻微下降;而小模型在推理任务上获得显著改进。本研究提出了一种可扩展且有效的语言模型创造力训练方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet the subjective nature of creativity and the limits of human judgment make training LLMs for creativity especially challenging. As a solution, we train LLMs on Codenames, a word-association game that exercises the two central axes of creativity, divergent and convergent thinking, while yielding objectively verifiable outcomes. This verifiability lets us bypass human judgment and train with Reinforcement Learning with Verifiable Rewards (RLVR). We train Qwen3-1.7B, 4B, and 8B models and evaluate them on ten creativity and four reasoning benchmarks. We find that the precision-diversity trade-off is scale-dependent: the 8B model prioritizes creativity over precision, while the 1.7B and 4B models gain reasoning precision at the cost of creativity. Concretely, the 8B model shows modest but consistent creativity gains (8 of 10 benchmarks) with only minor reasoning degradation, whereas the smaller models achieve substantial gains on reasoning tasks. Our study presents a scalable and effective solution to train LLMs for creativity.

语言模型创造力强化学习游戏训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。