arXiv:2604.10856cs.ROcs.AI2026-04被引 6

破解自动驾驶开环到闭环的性能下滑难题,提出实时自适应新方法

BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving

  • 发现开环预训练策略存在观测域偏移和目标错配两大根源问题
  • 提出测试时自适应框架,显著降低状态动作偏差并提升时序一致性
  • 揭示现有开环评估协议的盲区,适用于追求真实部署表现的自动驾驶研究

开环(OL)到闭环(CL)性能差距存在于开环评估表现优异的策略在闭环部署中效果不佳的现象。本文揭示了这一系统性失败的根本原因,并提出实用解决方案。具体而言,我们证明了开环策略存在观测域偏移和目标错配问题。尽管前者可通过适配技术部分修复,但后者导致其无法建模复杂反应行为,构成主要的OL-CL差距。我们发现多种开环策略均学习到一种忽略闭环仿真反应特性和减少累积误差所需时间感知的有偏Q值估计器。为此,我们提出一种测试时自适应(TTA)框架,用于校准观测偏移、降低状态-动作偏差并强制时序一致性。大量实验表明,该框架有效缓解规划偏差,且展现出优于基线的扩展动态特性。此外,我们的分析揭示标准开环评估协议存在盲点,无法反映闭环部署的真实情况。

原文摘要 · Abstract (English)

Open-loop (OL) to closed-loop (CL) gap (OL-CL gap) exists when OL-pretrained policies scoring high in OL evaluations fail to transfer effectively in closed-loop (CL) deployment. In this paper, we unveil the root causes of this systemic failure and propose a practical remedy. Specifically, we demonstrate that OL policies suffer from Observational Domain Shift and Objective Mismatch. We show that while the former is largely recoverable with adaptation techniques, the latter creates a structural inability to model complex reactive behaviors, which forms the primary OL-CL gap. We find that a wide range of OL policies learn a biased Q-value estimator that neglects both the reactive nature of CL simulations and the temporal awareness needed to reduce compounding errors. To this end, we propose a Test-Time Adaptation (TTA) framework that calibrates observational shift, reduces state-action biases, and enforces temporal consistency. Extensive experiments show that TTA effectively mitigates planning biases and yields superior scaling dynamics than its baseline counterparts. Furthermore, our analysis highlights the existence of blind spots in standard OL evaluation protocols that fail to capture the realities of closed-loop deployment.

自动驾驶强化学习测试时适应闭环评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。