arXiv:2511.23193cs.ROcs.LG2025-11

提升自动驾驶车辆在感知干扰下的协同驾驶能力

Fault-Tolerant MARL for CAVs under Observation Perturbations for Highway On-Ramp Merging

  • 引入对抗性扰动生成器强化训练鲁棒性
  • 自诊断机制利用时空相关性修复错误感知
  • 适用于高速公路汇入场景的高可靠性协同控制

多智能体强化学习(MARL)在实现联网自动驾驶汽车(CAVs)协同驾驶方面具有巨大潜力,但其实际应用受限于对观测故障的不足容错能力。这类故障表现为车辆感知数据中的扰动,会显著降低基于MARL的驾驶系统性能。解决该问题面临两大挑战:一是生成能有效考验策略的对抗性扰动,二是使车辆具备缓解污染观测影响的能力。为此,我们提出一种容错型MARL方法,包含两个关键智能体:首先,共训练一个对抗性故障注入智能体,主动生成扰动以强化车辆策略;其次,设计一种新型容错车辆智能体,具备自诊断能力,利用车辆状态序列中的固有时空相关性检测故障并重建可信观测,从而保护策略免受误导输入影响。在模拟高速公路汇入场景的实验表明,该方法显著优于基线MARL方法,在多种观测故障模式下均达到接近无故障水平的安全性与效率。

原文摘要 · Abstract (English)

Multi-Agent Reinforcement Learning (MARL) holds significant promise for enabling cooperative driving among Connected and Automated Vehicles (CAVs). However, its practical application is hindered by a critical limitation, i.e., insufficient fault tolerance against observational faults. Such faults, which appear as perturbations in the vehicles' perceived data, can substantially compromise the performance of MARL-based driving systems. Addressing this problem presents two primary challenges. One is to generate adversarial perturbations that effectively stress the policy during training, and the other is to equip vehicles with the capability to mitigate the impact of corrupted observations. To overcome the challenges, we propose a fault-tolerant MARL method for cooperative on-ramp vehicles incorporating two key agents. First, an adversarial fault injection agent is co-trained to generate perturbations that actively challenge and harden the vehicle policies. Second, we design a novel fault-tolerant vehicle agent equipped with a self-diagnosis capability, which leverages the inherent spatio-temporal correlations in vehicle state sequences to detect faults and reconstruct credible observations, thereby shielding the policy from misleading inputs. Experiments in a simulated highway merging scenario demonstrate that our method significantly outperforms baseline MARL approaches, achieving near-fault-free levels of safety and efficiency under various observation fault patterns.

多智能体强化学习自动驾驶容错控制感知扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。