arXiv:2509.08221cs.RO2025-09综述被引 12

系统梳理100篇基于CARLA的强化学习自动驾驶研究,揭示主流方法与挑战。

A Comprehensive Review of Reinforcement Learning for Autonomous Driving in the CARLA Simulator

  • 按算法类型分类,超80%研究采用无模型方法如DQN、PPO、SAC
  • 总结常见状态、动作与奖励设计,涵盖传感器模态与控制抽象方式
  • 提炼关键评估指标与场景配置,指明稀疏奖励、仿真到现实迁移等核心难题

自动驾驶研究近年来广泛采用深度强化学习(RL)作为数据驱动决策的框架,但现有算法在CARLA仿真器中的应用、基准测试与评估方式仍缺乏清晰图景。本综述系统分析了约100篇经同行评审的论文,这些论文在开源CARLA仿真器中训练、测试或验证RL策略。首先按算法族(无模型、基于模型、分层、混合)分类并量化其使用比例,发现超过80%的研究仍依赖无模型方法,如DQN、PPO和SAC。其次,阐述不同研究中状态、动作与奖励设计的多样性,说明传感器模态(RGB、LiDAR、BEV、语义地图、carla运动学状态)、控制抽象(离散与连续)及奖励塑造策略的选择差异。同时,整合评估体系,列出常用指标(成功率、碰撞率、车道偏离度、驾驶评分)及所用城镇、场景与交通配置。持续存在的挑战包括稀疏奖励、仿真到真实世界的迁移、安全保证不足及行为多样性有限,并据此提出开放性研究问题;展望了基于模型的强化学习、元学习及更丰富的多智能体仿真等有前景方向。通过提供统一分类体系、定量统计与局限性批判性讨论,本综述旨在为新手提供参考,并为推进基于强化学习的自动驾驶向真实部署迈进指明路径。

原文摘要 · Abstract (English)

Autonomous-driving research has recently embraced deep Reinforcement Learning (RL) as a promising framework for data-driven decision making, yet a clear picture of how these algorithms are currently employed, benchmarked and evaluated is still missing. This survey fills that gap by systematically analysing around 100 peer-reviewed papers that train, test or validate RL policies inside the open-source CARLA simulator. We first categorize the literature by algorithmic family model-free, model-based, hierarchical, and hybrid and quantify their prevalence, highlighting that more than 80% of existing studies still rely on model-free methods such as DQN, PPO and SAC. Next, we explain the diverse state, action and reward formulations adopted across works, illustrating how choices of sensor modality (RGB, LiDAR, BEV, semantic maps, and carla kinematics states), control abstraction (discrete vs. continuous) and reward shaping are used across various literature. We also consolidate the evaluation landscape by listing the most common metrics (success rate, collision rate, lane deviation, driving score) and the towns, scenarios and traffic configurations used in CARLA benchmarks. Persistent challenges including sparse rewards, sim-to-real transfer, safety guarantees and limited behaviour diversity are distilled into a set of open research questions, and promising directions such as model-based RL, meta-learning and richer multi-agent simulations are outlined. By providing a unified taxonomy, quantitative statistics and a critical discussion of limitations, this review aims to serve both as a reference for newcomers and as a roadmap for advancing RL-based autonomous driving toward real-world deployment.

强化学习自动驾驶CARLA综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。