用偏好优化方法跨架构蒸馏大模型,效果优于传统方法。
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
- 将知识蒸馏转化为偏好优化任务,利用多样化推理路径
- 在5个数据集上均超越传统黑盒蒸馏基线
- 采用混合策略提升学生模型输出利用率,适合多架构适配
我们提出ORPO-Distill,一种通用的跨架构大模型蒸馏方法,将问题建模为偏好优化任务。与传统的思维链蒸馏不同,该方法通过多样化的推理轨迹传递知识。采用几率比偏好优化目标,对比教师与学生推理轨迹以实现更高效学习,并引入混合策略利用学生生成的输出,在多个学生模型和五个数据集上的实验表明,其性能持续优于离线与在线策略基准。
原文摘要 · Abstract (English)
We introduce ORPO-Distill, a general-purpose method for cross-architecture LLM distillation that formulates the problem as a preference optimization task. Unlike standard CoT distillation, the approach transfers knowledge through diverse reasoning traces. It employs an Odds-Ratio Preference Optimization objective that contrasts teacher and student traces for more effective learning, and adopts a mixed-policy strategy for utilizing student-generated outputs, outperforming both off- and on-policy alternatives. Experiments on five datasets and multiple student models show consistent improvements over conventional black-box KD baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。