arXiv:2606.14368cs.LGcs.CL2026-06

两个大模型互当老师,共同提升跨领域能力而不丢专长。

Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback

论文配图:Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
图 1 · 摘自论文原文
  • 双模型基于自身正确回答和对方反馈进行自我优化。
  • 在科学问答任务中,双方均实现跨领域性能提升且不削弱原有优势。
  • 通过智能判断反馈时机与锚定问题,提升互教效果,适合多领域训练场景。

我们研究多领域大模型训练,其中两个在不同领域各有专长的模型,通过相互提供在线策略反馈实现共同进化。不同于单向知识蒸馏或单模型微调,目标是实现双方在所有领域上的帕累托改进:每个模型在跨域表现提升的同时,不损失原有专长。为此,我们提出在线策略协同蒸馏(OPCoD),其中每个模型的自蒸馏过程基于自身正确轨迹及其同伴的反馈。为提升反馈有效性,OPCoD采用认知感知门控机制决定何时给予反馈,并使用反馈锚定技术将反馈与具体问题绑定。在科学问答任务上,OPCoD持续优于基线方法,在所有评估的领域组合与学生模型中均实现帕累托改进。

原文摘要 · Abstract (English)

We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillation or single-model fine-tuning, our goal is mutual Pareto improvement: each model improves across domains without losing its original strength. To this end, we propose On-Policy Co-Distillation (OPCoD), where each student's self-distillation is conditioned on its own correct rollout and feedback from its peer. To make feedback exchange effective, OPCoD uses cognizance-based gating to decide when to give feedback and feedback anchoring to ground feedback in the problem. On Science Q\&A tasks, OPCoD consistently outperforms baselines and achieves Pareto improvement across all evaluated domain pairs and students.

大模型互训知识蒸馏多领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。