arXiv:2507.16382cs.ROcs.AI2025-07中稿 · IROS 2025

用大模型生成动态奖励函数,让多智能体更快实现编队避障。

Application of LLM Guided Reinforcement Learning in Formation Control with Collision Avoidance

  • 用大语言模型设计可在线调整的奖励函数,指导智能体决策。
  • 在仿真和真实场景中均减少迭代次数,更快达到最优性能。
  • 适合需要高效协同的机器人编队、自动驾驶等动态环境应用。

多智能体系统(MAS)通过个体协作完成复杂任务,其中多智能体强化学习(MARL)尤为有效。然而,在编队控制与避障(FCCA)这类复杂目标下,设计能快速收敛的奖励函数仍具挑战。本文提出新框架:利用大语言模型(LLM)分析任务优先级和各智能体可观测信息,生成可基于评估结果动态调整的奖励函数,而非直接依赖原始奖励。该机制使多智能体系统在动态环境中更高效地同步实现编队与避障,显著减少迭代次数即可达到更高性能。实证研究在仿真与真实场景中验证了该方法的实用性与有效性。

原文摘要 · Abstract (English)

Multi-Agent Systems (MAS) excel at accomplishing complex objectives through the collaborative efforts of individual agents. Among the methodologies employed in MAS, Multi-Agent Reinforcement Learning (MARL) stands out as one of the most efficacious algorithms. However, when confronted with the complex objective of Formation Control with Collision Avoidance (FCCA): designing an effective reward function that facilitates swift convergence of the policy network to an optimal solution. In this paper, we introduce a novel framework that aims to overcome this challenge. By giving large language models (LLMs) on the prioritization of tasks and the observable information available to each agent, our framework generates reward functions that can be dynamically adjusted online based on evaluation outcomes by employing more advanced evaluation metrics rather than the rewards themselves. This mechanism enables the MAS to simultaneously achieve formation control and obstacle avoidance in dynamic environments with enhanced efficiency, requiring fewer iterations to reach superior performance levels. Our empirical studies, conducted in both simulation and real-world settings, validate the practicality and effectiveness of our proposed approach.

强化学习多智能体编队控制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。