两个大模型互当老师,共同提升跨领域能力而不丢专长。
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback

- 双模型基于自身正确回答和对方反馈进行自我优化。
- 在科学问答任务中,双方均实现跨领域性能提升且不削弱原有优势。
- 通过智能判断反馈时机与锚定问题,提升互教效果,适合多领域训练场景。
我们研究多领域大模型训练,其中两个在不同领域各有专长的模型,通过相互提供在线策略反馈实现共同进化。不同于单向知识蒸馏或单模型微调,目标是实现双方在所有领域上的帕累托改进:每个模型在跨域表现提升的同时,不损失原有专长。为此,我们提出在线策略协同蒸馏(OPCoD),其中每个模型的自蒸馏过程基于自身正确轨迹及其同伴的反馈。为提升反馈有效性,OPCoD采用认知感知门控机制决定何时给予反馈,并使用反馈锚定技术将反馈与具体问题绑定。在科学问答任务上,OPCoD持续优于基线方法,在所有评估的领域组合与学生模型中均实现帕累托改进。
原文摘要 · Abstract (English)
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillation or single-model fine-tuning, our goal is mutual Pareto improvement: each model improves across domains without losing its original strength. To this end, we propose On-Policy Co-Distillation (OPCoD), where each student's self-distillation is conditioned on its own correct rollout and feedback from its peer. To make feedback exchange effective, OPCoD uses cognizance-based gating to decide when to give feedback and feedback anchoring to ground feedback in the problem. On Science Q\&A tasks, OPCoD consistently outperforms baselines and achieves Pareto improvement across all evaluated domain pairs and students.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。