TREX用多智能体自动完成大模型训练全流程,提升效率与效果。
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration

- 构建研究者与执行者协作的多智能体系统,驱动训练流程自动化。
- 在10个真实场景任务上持续优化模型性能,实现高效探索与迭代。
- 适合想自动化训练大模型的研究者,尤其关注端到端流程优化。
尽管大语言模型已能完成孤立的科研任务,但自动化复杂、现实世界的工作流(如大模型训练)仍面临重大挑战。本文提出TREX,一个用于自动化大模型训练全生命周期的多智能体系统。通过协调核心模块——研究者与执行者之间的协作,系统可无缝完成需求分析、开放域文献与数据调研、训练策略制定、数据配方准备,以及模型训练与评估。多轮实验过程被建模为搜索树,使系统能高效规划探索路径、复用历史结果,并从迭代试验中提炼高层洞察。为评估自动化大模型训练能力,我们构建了FT-Bench基准,包含10个源自真实场景的任务,涵盖基础模型能力优化到领域特定任务性能提升。实验结果表明,TREX智能体在目标任务上持续优化模型表现。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have empowered AI research agents to perform isolated scientific tasks, automating complex, real-world workflows, such as LLM training, remains a significant challenge. In this paper, we introduce TREX, a multi-agent system that automates the entire LLM training life-cycle. By orchestrating collaboration between two core modules-the Researcher and the Executor-the system seamlessly performs requirement analysis, open-domain literature and data research, formulation of training strategies, preparation of data recipes, and model training and evaluation. The multi-round experimental process is modeled as a search tree, enabling the system to efficiently plan exploration paths, reuse historical results, and distill high-level insights from iterative trials. To evaluate the capability of automated LLM training, we construct FT-Bench, a benchmark comprising 10 tasks derived from real-world scenarios, ranging from optimizing fundamental model capabilities to enhancing performance on domain-specific tasks. Experimental results demonstrate that the TREX agent consistently optimizes model performance on target tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。