arXiv:2605.10781cs.LGcs.CL2026-05被引 9

让大模型在成功时自主探索,突破传统引导束缚。

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

论文配图:Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
图 1 · 摘自论文原文
  • 反向解读教师信号,识别学生自主推理的正确路径
  • 在成功推理中强化学生原创选择,提升性能表现
  • 适合追求模型自主思考能力的研究者与开发者

自蒸馏已成为后训练大语言模型的强大框架,其中教师基于额外信息指导无此信息的学生,二者来自同一模型。然而,当学生表现良好时,该机制反而覆盖学生选择,抑制其自身推理。为此,我们提出逆向解读原始自蒸馏信号:当学生沿教师无法预测的路径取得成功时,这些标记反映其自主推理。基于此,我们提出RLRT(RLVR中教师信号反转),通过增强正确推演中的此类标记来改进GRPO。这被视为一种新形式的探索——非均匀多样性,而是以学生成功为基础的价值探索。在基础、指令微调及思维微调的Qwen3检查点上,RLRT显著优于自蒸馏与探索基线,确立信息不对称作为RLVR的新原则设计维度。

原文摘要 · Abstract (English)

Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.

自蒸馏强化学习推理探索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。