arXiv:2602.18097cs.ROcs.LG2026-02

融合安全分析与强化学习,让自动驾驶车更安全地与骑车人互动。

Interacting safely with cyclists using Hamilton-Jacobi reachability and reinforcement learning

  • 用哈密顿-雅可比可达性分析构建安全度量
  • 将安全指标作为奖励信号优化导航策略
  • 模拟测试显示接近人类驾驶行为,优于现有方法

本文提出一种框架,使自动驾驶车辆在与骑车人交互时兼顾安全与效率。该方法结合哈密顿-雅可比可达性分析与深度Q-learning,通过求解时变哈密顿-雅可比-贝尔曼不等式获得每个系统状态的安全度量,作为结构化奖励信号嵌入强化学习中。同时建模骑车人的潜在响应,通过扰动输入反映人类舒适度与行为适应性。在仿真中评估了该框架,并与人类驾驶行为及一种现有先进方法进行对比。

原文摘要 · Abstract (English)

In this paper, we present a framework for enabling autonomous vehicles to interact with cyclists in a manner that balances safety and optimality. The approach integrates Hamilton-Jacobi reachability analysis with deep Q-learning to jointly address safety guarantees and time-efficient navigation. A value function is computed as the solution to a time-dependent Hamilton-Jacobi-Bellman inequality, providing a quantitative measure of safety for each system state. This safety metric is incorporated as a structured reward signal within a reinforcement learning framework. The method further models the cyclist's latent response to the vehicle, allowing disturbance inputs to reflect human comfort and behavioral adaptation. The proposed framework is evaluated through simulation and comparison with human driving behavior and an existing state-of-the-art method.

自动驾驶安全交互强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。