arXiv:2602.02709cs.AI2026-02

多智能体框架ATLAS通过动态参考模型提升强化学习自进化能力。

ATLAS: A Multi-LLM Training Framework for EvoDPO with Adaptive Reference Evolution

  • 引入可自适应更新的参考策略,替代固定参考模型
  • 在多种复杂任务中实现长期评估下的稳定性能提升
  • 适合研究智能体自进化与动态偏好优化的学者

近期多大语言模型(LLM)智能体系统展现出自动解题的潜力,但大多依赖冻结的智能体或静态微调流程。为此,我们提出ATLAS(面向智能体自进化的自适应任务分配学习框架),一个由专业化元智能体协同训练并优化活跃智能体以达成特定领域策略的多智能体框架。迭代偏好学习中的核心挑战在于对固定参考模型的依赖,常导致更新过于保守或训练停滞。为克服此问题,框架采用演进式直接偏好优化(EvoDPO),利用检查智能体基于持续训练遥测数据,实现自适应、代理KL门控的参考策略更新。我们在多样且具挑战性的环境(包括非平稳上下文老虎机、偏微分方程求解、组合优化任务如旅行商问题与装箱问题)中评估该框架。相较于固定参考、自适应参考及外部自动化发现基线,结果表明ATLAS结合支持者驱动探索与EvoDPO驱动的稳定性,在长周期评估下显著提升了自改进能力。

原文摘要 · Abstract (English)

Recent multi-LLM agent systems have shown promising capabilities for automated problem-solving, yet they predominantly rely on frozen agents or static fine-tuning pipelines. To address this limitation, our primary contribution is ATLAS (Adaptive Task-distributed Learning for Agentic Self-evolution), a multi-agent framework where specialized meta-agents collaboratively train and refine an active agent toward a domain-specific policy. A core challenge in iterative preference learning within these pipelines is the reliance on fixed reference models, which typically leads to overly conservative updates or training stagnation. To overcome this, the framework's algorithmic engine utilizes Evolving Direct Preference Optimization (EvoDPO). EvoDPO employs an inspection agent to perform adaptive, proxy-KL gated reference policy updates based on continuous training telemetry. We evaluate this full framework across a diverse set of challenging environments-including non-stationary contextual bandits, partial differential equations (PINNs), and combinatorial optimization tasks (TSP, Bin Packing). Through comparison against fixed-reference, adaptive-reference, and external automated-discovery baselines, our results suggest that ATLAS combines supporter-driven exploration with EvoDPO-driven stability to improve long-horizon evaluator-driven self-improvement.

多智能体自进化偏好优化动态参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。