arXiv:2509.25100cs.LGcs.AI2025-09中稿 · NeurIPS被引 2

用偏好优化方法跨架构蒸馏大模型,效果优于传统方法。

ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation

  • 将知识蒸馏转化为偏好优化任务,利用多样化推理路径
  • 在5个数据集上均超越传统黑盒蒸馏基线
  • 采用混合策略提升学生模型输出利用率,适合多架构适配

我们提出ORPO-Distill,一种通用的跨架构大模型蒸馏方法,将问题建模为偏好优化任务。与传统的思维链蒸馏不同,该方法通过多样化的推理轨迹传递知识。采用几率比偏好优化目标,对比教师与学生推理轨迹以实现更高效学习,并引入混合策略利用学生生成的输出,在多个学生模型和五个数据集上的实验表明,其性能持续优于离线与在线策略基准。

原文摘要 · Abstract (English)

We introduce ORPO-Distill, a general-purpose method for cross-architecture LLM distillation that formulates the problem as a preference optimization task. Unlike standard CoT distillation, the approach transfers knowledge through diverse reasoning traces. It employs an Odds-Ratio Preference Optimization objective that contrasts teacher and student traces for more effective learning, and adopts a mixed-policy strategy for utilizing student-generated outputs, outperforming both off- and on-policy alternatives. Experiments on five datasets and multiple student models show consistent improvements over conventional black-box KD baselines.

大模型蒸馏偏好优化跨架构推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。