通过多轮辩论与树状偏好优化,让小模型高效学习大模型的推理能力。
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
- 小模型与大模型多轮辩论,提取纠错策略等可操作反馈。
- 用树状结构组织辩论日志,提升训练效率与性能。
- 在多个任务上显著提升小模型准确率与泛化能力,适合资源受限场景。
大型语言模型在知识密集型和复杂推理任务中持续刷新标准,但其高计算开销限制了广泛应用。尽管将大模型压缩为小模型是可持续的解决方案,但现有技术如静态知识蒸馏、依赖大量资源的基于人类反馈的强化学习或有限的自我反思,难以带来显著且持久的性能提升。本文提出一种新的辩论与反思(D&R)框架,通过小模型与强教师模型进行多轮辩论,获取可操作的反馈(如错误分析、纠正策略),指导学生模型优化。进一步提出树状直接偏好优化(T-DPO),将辩论日志以层级结构组织,实现高效训练。在多个自然语言处理基准上的实证评估表明,该方法显著提升了小模型的准确率、鲁棒性和泛化能力,大幅优于传统基线。
原文摘要 · Abstract (English)
Large Language Models (LLMs) continue to set new standards in knowledge-intensive and complex reasoning tasks, yet their high computational demands limit widespread adoption. While distilling large models into smaller ones offers a sustainable solution, current techniques--such as static knowledge distillation, resource-intensive reinforcement learning from human feedback, or limited self-reflection--struggle to yield substantial and lasting performance gains. In this paper, we present a novel Debate and Reflect (D&R) framework that orchestrates multi-turn debates between smaller models and stronger teacher models, eliciting actionable feedback (e.g., error analysis, corrective strategies) to guide student models. Further, we introduce Tree-structured Direct Preference Optimization (T-DPO) to efficiently leverage these debate logs, organizing interactions into a hierarchical format for effective training. Empirical evaluations across diverse NLP benchmarks demonstrate that our approach significantly improves smaller-model accuracy, robustness, and generalization, outperforming conventional baselines by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。