arXiv:2511.15694cs.LG2025-11中稿 · NeurIPS被引 2

量化影响大模型推理强化学习,感知训练更优。

The Impact of Quantization on Large Reasoning Model Reinforcement Learning

  • 在强化学习中对比后训练量化与感知训练量化效果
  • 后训练量化在数学推理上表现优于感知训练量化
  • 适合关注大模型量化部署与强化学习结合的研究者

大规模强化学习(RL)如今可在无需监督微调的情况下实现强大推理能力。尽管后训练量化(PTQ)和量化感知训练(QAT)在微调场景中已得到充分研究,但量化对大型推理模型(LRMs)在强化学习中的影响仍不明确。为解答该问题,我们开展了系统性实验,发现经强化学习后的量化模型与量化感知强化学习优化模型在数学基准测试中存在显著性能差距。结果表明,量化感知强化学习训练反而对学习过程产生负面影响,而PTQ与QLoRA则带来更高性能。

原文摘要 · Abstract (English)

Strong reasoning capabilities can now be achieved by large-scale reinforcement learning (RL) without any supervised fine-tuning. Although post-training quantization (PTQ) and quantization-aware training (QAT) are well studied in the context of fine-tuning, how quantization impacts RL in large reasoning models (LRMs) remains an open question. To answer this question, we conducted systematic experiments and discovered a significant gap in reasoning performance on mathematical benchmarks between post-RL quantized models and their quantization-aware RL optimized counterparts. Our findings suggest that quantization-aware RL training negatively impacted the learning process, whereas PTQ and QLoRA led to greater performance.

量化强化学习大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。