arXiv:2603.18363cs.CLcs.AI2026-03中稿 · ICML被引 1

通过分布匹配让大模型同时具备逻辑推理与创造力

PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

  • 将无监督微调转化为分布匹配问题,用轨迹平衡目标消除生成长度偏差
  • 调节参数α可控制模型在逻辑严谨与创意表达间的平衡,提升综合性能
  • 适合追求模型能力多样性的研究者,尤其在创意生成任务中优势显著

无监督强化学习从内部反馈(RLIF)已成为激发大语言模型(LLMs)潜在能力的新范式,无需外部监督。然而现有方法依赖启发式内在奖励,缺乏明确的理论优化目标,易产生退化偏差。本文提出PowerFlow,将无监督微调重新建模为分布匹配问题。通过将GFlowNet作为非归一化密度的近似变分采样器,提出长度感知的轨迹平衡目标,显式消除自回归生成中的结构长度偏差。通过瞄准α-幂分布,PowerFlow实现对LLMs双重特性的定向激发:当α>1时聚焦逻辑推理,α<1时释放表达创造力。大量实验表明,PowerFlow持续优于现有RLIF方法,达到甚至超越有监督的GRPO表现。此外,通过缓解对齐模型的过度锐化,本方法在创造性任务中同时提升多样性与质量,推动了帕累托前沿的前移。

原文摘要 · Abstract (English)

Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting $α$-power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution ($α> 1$) to intensify logical reasoning, or flattening it ($α< 1$) to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.

大模型训练分布匹配创造力增强强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。