arXiv:2607.13348cs.RO2026-07

提出分层框架,让自动驾驶赛车在保证安全前提下更灵活超车。

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

论文配图:Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control
图 1 · 摘自论文原文
  • 高阶优化选超车路径,低阶控制用学习自适应安全约束
  • 自适应调节安全参数,跨赛道表现更稳定,成功率提升
  • 适合需兼顾安全与速度的复杂场景,如自动驾驶赛车

自动驾驶赛车超车需在非线性车辆动力学和实时约束下平衡性能与安全。模型预测控制(MPC)结合控制屏障函数(CBF)可保证安全集的前向不变性,但常见的固定衰减离散时间CBF方法在交互式竞速中过于保守,限制超车性能并需人工调参。本文提出分层超车框架,将决策与安全轨迹控制分离:高阶混合整数二次规划(MIQP)求解可行超车拓扑,低阶非线性弗雷内坐标系MPC通过嵌入离散时间CBF约束实现动力学与安全控制。该分解将组合复杂度与连续轨迹优化解耦。为缓解固定衰减约束的敏感性,引入强化学习策略在线调整CBF衰减参数,实现上下文感知的安全裕度调节,不直接控制输入。仿真与缩比硬件实验表明,单一固定衰减参数无法在各赛道保持优异表现,而自适应策略在无每赛道调参情况下获得最高综合成功率,持续保持良好安全-性能权衡,提升环境变化鲁棒性,并在正常运行中满足安全约束。

原文摘要 · Abstract (English)

Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Predictive Control (MPC) combined with Control Barrier Functions (CBFs) provides a principled mechanism for certifying forward invariance of a safe set. However, commonly used fixed-decay discrete-time CBF formulations can become overly conservative in interactive racing scenarios, limiting overtaking performance and requiring manual tuning across track conditions. This paper proposes a hierarchical overtaking framework that explicitly separates maneuver-level decision making from safety-certified trajectory control, reducing conservatism while preserving safety. A high-level Mixed-Integer Quadratic Program (MIQP) resolves the combinatorial passing-side selection problem by selecting a feasible overtaking topology, while a nonlinear Frenet-frame MPC enforces vehicle dynamics and safety through embedded discrete-time CBF constraints. This decomposition isolates the combinatorial complexity of maneuver selection from the continuous trajectory optimization. To further mitigate the sensitivity of fixed-decay barrier constraints, a reinforcement learning policy adapts the discrete-time CBF decay parameter online, enabling context-dependent modulation of safety margins without directly controlling vehicle inputs. Simulation and scaled-hardware experiments show that no single fixed decay parameter achieves uniformly strong performance across tracks, whereas the adaptive strategy attains the highest aggregate success rate and consistently strong safety--performance trade-offs without per-track tuning, improving robustness to environment variation while maintaining safety constraint satisfaction in nominal operation.

自动驾驶安全控制分层优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。