让大模型像人一样进化,从依赖指导到自我成长。
METEOR: Evolutionary Journey of Large Language Models from Guidance to Self-Growth
- 分三阶段训练:数据蒸馏、迭代优化、自我演化
- 在多个领域任务中显著提升准确率与可靠性
- 适合希望实现模型自主进化的研究者和开发者
模型进化使模型能从反馈中学习,不断精炼经验并更新技能,实现从无领域知识到成为领域专家的转变。然而,目前尚无统一有效的机制来引导这一演化过程。为此,我们提出Meteor方法,包含三个训练阶段:弱到强的数据蒸馏、迭代训练和自演化策略。每个阶段均最大化模型内在领域能力,使其能够自主完善领域知识并提升性能。实验表明,该方法在特定领域任务中显著提升了准确性、完整性、相关性、连贯性和可靠性。
原文摘要 · Abstract (English)
Model evolution enables learning from feedback to refine experiences and update skills, transforming models from having no domain knowledge to becoming domain experts. However, there is currently no unified and effective method for guiding this evolutionary process. To address this gap, we propose the Meteor method, which includes three training phases: weak-to-strong data distillation, iterative training, and self-evolution strategies. Each phase maximizes the model's inherent domain capabilities, allowing it to autonomously refine its domain knowledge and enhance performance. Experiments demonstrate that our approach significantly improves accuracy, completeness, relevance, coherence, and reliability across domain-specific tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。