提出新框架平衡大模型推理与创造力,防止思维路径单一化
The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
- 用概率分布视角统一多种推理训练方法,揭示其本质共性
- 证明纠错导向训练会导致思维多样性衰减,影响创造性解题
- 给出可操作方案,让模型既准确又保持思维多样性
当前最先进的大语言模型推理流程依赖自举式推理循环:采样多样化的思维链并强化得分最高的路径,主要优化正确性。我们分析发现,这种设计对模型推理路径分布的坍缩高度敏感,导致语义熵下降,削弱创造性问题解决能力。为此,我们提出分布式创造性推理(DCR),一种统一的变分目标,将训练视为在解题轨迹概率测度上的梯度流。STaR、GRPO、DPO、熵奖励等方法均为该损失的特例。该框架带来三大核心结果:(i) 多样性衰减定理,描述了基于正确性的目标如何导致STaR、GRPO和DPO产生不同模式的多样性衰减;(ii) 设计能保证收敛到稳定且多样化的策略,有效防止分布坍缩;(iii) 实用的可操作方法,可在实践中实现。DCR因此提供了首个兼具准确性和创造性的大模型推理原则性方案。
原文摘要 · Abstract (English)
State-of-the-art large language model (LLM) pipelines rely on bootstrapped reasoning loops: sampling diverse chains of thought and reinforcing the highest-scoring ones, mainly optimizing correctness. We analyze how this design choice is sensitive to the collapse of the model's distribution over reasoning paths, slashing semantic entropy and undermining creative problem-solving. To analyze this failure, we introduce Distributional Creative Reasoning (DCR), a unified variational objective that casts training as gradient flow through probability measures on solution traces. STaR, GRPO, and DPO, as well as entropy bonuses, and other methods, all constitute special cases of the same loss. The framework delivers three core results: (i) the diversity decay theorem, describing how correctness-based objectives lead to distinct modes of diversity decay for STaR, GRPO, and DPO; (ii) designs that ensure convergence to a stable and diverse policy, effectively preventing collapse; and (iii) simple, actionable recipes to achieve this in practice. DCR thus offers the first principled recipe for LLMs that remain both correct and creative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。