arXiv:2602.04620cs.LG2026-02被引 1

提出QUATRO,让大模型微调更稳定可控。

QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning

  • 基于可解释的约束优化,动态调节策略更新幅度。
  • 在数学推理任务上,支持高学习率和策略过时场景仍稳定训练。
  • 适合追求训练稳定性和可控性的大模型调优研究者。

基于强化学习的GRPO类大模型微调方法近期受到关注。然而,依赖启发式信任区域近似的方法容易导致优化行为脆弱,因为全局重要性比率裁剪和组内归一化无法有效调控超出裁剪范围样本的影响。我们提出查询自适应信任区域策略优化(QUATRO),通过原理化的优化直接施加信任区域约束。这带来了清晰可解释的目标函数,实现对策略更新的显式控制与稳定的熵可控优化,且稳定项由精确的信任区域公式自然产生。在多种数学推理基准上的实证验证表明,QUATRO在策略过时加剧和采用激进学习率的情况下仍能保持稳定训练,并全程维持良好控制的熵值。

原文摘要 · Abstract (English)

GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate samples whose importance ratios fall outside the clipping range. We propose Query-Adaptive Trust-Region policy Optimization (QUATRO), which directly enforces trust-region constraints through a principled optimization. This yields a clear and interpretable objective that enables explicit control over policy updates and stable, entropy-controlled optimization, with a stabilizer terms arising intrinsically from the exact trust-region formulation. Empirically verified on diverse mathematical reasoning benchmarks, QUATRO shows stable training under increased policy staleness and aggressive learning rates, maintaining well-controlled entropy throughout training.

大模型微调强化学习信任区域稳定训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。