arXiv:2501.05707cs.CLcs.AI2025-01ICLR被引 97

让多个AI模型互动生成数据,实现持续自我进化。

Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

  • 用多智能体互动生成多样化训练数据,各模型独立微调。
  • 在推理任务上表现优于单智能体方法,可支持更多轮自进化。
  • 适合需要持续优化的复杂推理场景,如AI研究与开发。

大型语言模型(LLMs)近年来取得显著进展,但其性能受限于原始训练数据。为突破这一限制,近期研究探索利用大模型生成合成数据以实现自主自我改进。然而,持续的自我改进会逐渐进入收益递减阶段。本文提出一种互补性自我改进方法:对由多个语言模型组成的多智能体社会进行微调。所有模型均从同一基础模型出发,通过在模型间多智能体交互中生成的数据独立更新,实现各自专业化。通过为每个模型提供独立的数据集,该方法实现了模型间的分工与整体多样性提升。结果表明,系统能长期保持多样化的推理路径,并在比单智能体方法更多的微调轮次中实现自主进化。我们在一系列推理任务中定量验证了该方法的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. However, successive steps of self-improvement can reach a point of diminishing returns. In this work, we propose a complementary approach towards self-improvement where finetuning is applied to a multiagent society of language models. A group of language models, all starting from the same base model, are independently specialized by updating each one using data generated through multiagent interactions among the models. By training each model on independent sets of data, we illustrate how this approach enables specialization across models and diversification over the set of models. As a result, our overall system is able to preserve diverse reasoning chains and autonomously improve over many more rounds of fine-tuning than single-agent self-improvement methods. We quantitatively illustrate the efficacy of the approach across a wide suite of reasoning tasks.

多智能体自进化推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。