arXiv:2608.01589cs.AI2026-08被引 1

用问题空间引导的自蒸馏方法,让模型更高效地学习解题结构。

Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning

论文配图:Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning
图 1 · 摘自论文原文
  • 用问题状态和路径替代完整解法作为教师信号
  • 在3个数学推理数据集上达到最高准确率(1.7B~8B模型)
  • 适合需要提升推理能力的模型优化场景

在策略自蒸馏(OPSD)中,通过参考解法提供的特权信息监督仅观察问题的学生模型,可提升推理能力。然而,教师给出的逐标记目标可能依赖推理时不可用的特定解法信息。本文提出问题空间引导的OPSD(PS-OPSD),以初始状态、目标条件、约束和选定的状态转移路径作为轨迹引导,取代完整解法。学生模型的推理过程与蒸馏目标保持不变。在三个数学推理基准上,模型规模从1.7B到8B,PS-OPSD在仅使用问题输入的情况下实现了最高的综合准确率。受控实验表明,引导相关性和路径一致性对性能提升有显著贡献,凸显了特权信息表征方式在OPSD中的关键作用。

原文摘要 · Abstract (English)

On-policy self-distillation (OPSD) improves reasoning by using a privileged view of a model conditioned on reference solutions to supervise a student view that observes only the question. However, the teacher-provided token-level targets may depend on reference-specific information unavailable at inference time. We propose Problem-Space-Guided OPSD (PS-OPSD), which replaces the complete solution with trajectory-grounded guidance describing the initial state, goal conditions, constraints, and a selected state-transition path. The student rollout and OPSD objective remain unchanged. Across three mathematical reasoning benchmarks and model scales ranging from 1.7B to 8B, PS-OPSD achieves the highest aggregate question-only accuracy among the compared methods. Controlled experiments further indicate that guidance relevance and path coherence contribute to these gains, highlighting the representation of privileged information as an important design choice in OPSD.

自蒸馏推理增强模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。