arXiv:2604.09499cs.RO2026-04被引 1

用物理启发的强化学习实现无地图赛车,效率高且真实场景表现优于人类。

Physics-Informed Reinforcement Learning of Spatial Density Velocity Potentials for Map-Free Racing

论文配图:Physics-Informed Reinforcement Learning of Spatial Density Velocity Potentials for Map-Free Racing
图 1 · 摘自论文原文
  • 通过深度图谱特征与物理奖励结合,端到端学习车辆动力学控制策略。
  • 在未见过的赛道上比人类示范快12%,计算量不足传统方法1%。
  • 适合需要高效实时控制的无人赛车系统,尤其关注泛化与硬件部署。

无地图自主赛车是嵌入式机器人的一项重大挑战,要求基于瞬时传感器数据,在加速度和轮胎摩擦极限下进行运动规划。为实现对多种赛道配置的分布外(OOD)泛化,机器学习被用于编码传感器数据与车辆执行之间的数学关系,实现端到端控制并隐含定位功能。行为克隆(BC)受限于人类反应速度,而深度强化学习(DRL)虽在仿真中表现更优,但需大量碰撞训练,难以迁移到真实世界,导致硬件上执行不稳定。本文提出一种新DRL方法:利用深度测量的谱分布参数化非线性车辆动力学,并设计非几何、物理启发的奖励函数,使人工神经网络(ANN)能推断时间最优及超车控制策略,计算量小于BC和基于模型的DRL的1%。通过物理引擎感知奖励与隐式价值域截断替代显式碰撞惩罚,消除从仿真到现实的迁移障碍和方差带来的保守性。该策略在按比例缩放的真实硬件上,于未知赛道上比人类示范快12%,并充分激活摩擦圆,轮胎动力学逼近实测的Pacejka模型。系统辨识揭示功能分岔:第一层将空间观测压缩以提取更高分辨率的弯道顶点特征,第二层编码非线性动力学。

原文摘要 · Abstract (English)

Autonomous racing without prebuilt maps is a grand challenge for embedded robotics that requires kinodynamic planning from instantaneous sensor data at the acceleration and tire friction limits. Out-Of-Distribution (OOD) generalization to various racetrack configurations utilizes Machine Learning (ML) to encode the mathematical relation between sensor data and vehicle actuation for end-to-end control, with implicit localization. These comprise Behavioral Cloning (BC) that is capped to human reaction times and Deep Reinforcement Learning (DRL) which requires large-scale collisions for comprehensive training that can be infeasible without simulation but is arduous to transfer to reality, thus exhibiting greater performance than BC in simulation, but actuation instability on hardware. This paper presents a DRL method that parameterizes nonlinear vehicle dynamics from the spectral distribution of depth measurements with a non-geometric, physics-informed reward, to infer vehicle time-optimal and overtaking racing controls with an Artificial Neural Network (ANN) that utilizes less than 1% of the computation of BC and model-based DRL. Slaloming from simulation to reality transfer and variance-induced conservatism are eliminated with the combination of a physics engine exploit-aware reward and the replacement of an explicit collision penalty with an implicit truncation of the value horizon. The policy outperforms human demonstrations by 12% in OOD tracks on proportionally scaled hardware, by maximizing the friction circle with tire dynamics that resemble an empirical Pacejka tire model. System identification illuminates a functional bifurcation where the first layer compresses spatial observations to extract digitized track features with higher resolution in corner apexes, and the second encodes nonlinear dynamics.

强化学习无人赛车物理启发端到端控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。