arXiv:2501.05057cs.ROcs.AI2025-01被引 3

用大模型自动设计驾驶策略训练流程,提升效率与泛化能力。

LearningFlow: Automated Policy Learning Workflow for Urban Driving with Large Language Models

  • 多大模型代理协作生成训练课程与奖励函数
  • 在CARLA仿真中实现更高样本效率和更强泛化性能
  • 适合自动驾驶研发团队快速迭代策略算法

强化学习在自动驾驶领域展现出巨大潜力,但奖励函数人工设计繁琐、复杂环境样本效率低等问题仍制约发展。为此,本文提出LearningFlow,一种面向城市驾驶的自动化策略学习工作流。该框架通过多个大语言模型代理协同,在强化学习训练过程中自动完成课程序列生成与奖励函数设计。每个环节均由分析代理评估训练进展并提供反馈,支持生成代理动态优化内容。实验在高保真CARLA模拟器中开展,对比多种现有方法,结果表明LearningFlow能有效生成高质量奖励与课程,在多样化驾驶任务中表现优异,具备出色的鲁棒泛化能力,并可适配不同强化学习算法。

原文摘要 · Abstract (English)

Recent advancements in reinforcement learning (RL) demonstrate the significant potential in autonomous driving. Despite this promise, challenges such as the manual design of reward functions and low sample efficiency in complex environments continue to impede the development of safe and effective driving policies. To tackle these issues, we introduce LearningFlow, an innovative automated policy learning workflow tailored to urban driving. This framework leverages the collaboration of multiple large language model (LLM) agents throughout the RL training process. LearningFlow includes a curriculum sequence generation process and a reward generation process, which work in tandem to guide the RL policy by generating tailored training curricula and reward functions. Particularly, each process is supported by an analysis agent that evaluates training progress and provides critical insights to the generation agent. Through the collaborative efforts of these LLM agents, LearningFlow automates policy learning across a series of complex driving tasks, and it significantly reduces the reliance on manual reward function design while enhancing sample efficiency. Comprehensive experiments are conducted in the high-fidelity CARLA simulator, along with comparisons with other existing methods, to demonstrate the efficacy of our proposed approach. The results demonstrate that LearningFlow excels in generating rewards and curricula. It also achieves superior performance and robust generalization across various driving tasks, as well as commendable adaptation to different RL algorithms.

自动驾驶强化学习大模型智能驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。