arXiv:2504.09021cs.LG2025-04中稿 · ICRA被引 7

首个基于视觉的赛车智能体,无需定位即可达到冠军级表现

A Champion-level Vision-based Reinforcement Learning Agent for Competitive Racing in Gran Turismo 7

  • 用自车视角摄像头+车载传感器实现纯视觉决策
  • 在GT7中持续击败内置电脑玩家,性能达职业水平
  • 适合对真实自动驾驶感知有需求的研究者

深度强化学习已在高保真模拟器Gran Turismo 7(GT7)中实现超人级赛车表现。传统方法依赖外部仪器提供的全局信息(如车辆与对手的精确位置),限制了现实应用。为此,本文提出一个仅依赖自车视角摄像头和车载传感器数据的视觉自主赛车智能体,推理时无需精确定位。该智能体采用非对称演员-评论家框架:演员使用含自车传感器数据的循环神经网络,记忆赛道布局与对手位置;评论家在训练阶段可访问全局特征。在GT7中的评估显示,该智能体持续优于游戏内置驾驶员。据我们所知,这是首个在竞技赛况下展现冠军级表现的视觉自主赛车智能体。

原文摘要 · Abstract (English)

Deep reinforcement learning has achieved superhuman racing performance in high-fidelity simulators like Gran Turismo 7 (GT7). It typically utilizes global features that require instrumentation external to a car, such as precise localization of agents and opponents, limiting real-world applicability. To address this limitation, we introduce a vision-based autonomous racing agent that relies solely on ego-centric camera views and onboard sensor data, eliminating the need for precise localization during inference. This agent employs an asymmetric actor-critic framework: the actor uses a recurrent neural network with the sensor data local to the car to retain track layouts and opponent positions, while the critic accesses the global features during training. Evaluated in GT7, our agent consistently outperforms GT7's built-drivers. To our knowledge, this work presents the first vision-based autonomous racing agent to demonstrate champion-level performance in competitive racing scenarios.

强化学习视觉导航赛车仿真端到端控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。