arXiv:2608.29772cs.RO2026-08

让自动驾驶系统学会识别自身不足并主动求助,持续改进安全表现。

Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving

论文配图:Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving
图 1 · 摘自论文原文
  • 构建自感知机制,通过恐惧与好奇信号判断何时需专家介入。
  • 在多种仿真与真实场景中,安全违规减少,性能保持稳定优于基线。
  • 适合追求高可靠性、可持续进化的自动驾驶系统研发团队使用。

基于学习的自动驾驶系统在熟悉环境中表现可靠,但罕见分布偏移和长尾事件仍是突发故障的主要原因。核心局限在于多数智能体仅被动积累经验,缺乏评估自身能力不足、及时寻求协助并从高危事件中针对性提升的能力。本文提出自感知引导探索(SAGE)框架,用于训练后的适应性优化。SAGE 学习一个预测世界模型,生成两个在线内在信号:恐惧信号衡量短时预测风险与模型不确定性;好奇信号通过预测误差衡量新颖性。好奇心动态校准恐惧的干预阈值,使代理能根据上下文调节风险控制。当预测恐惧超过该自适应阈值时,代理将控制权移交专家或备用策略,并利用接管轨迹进行聚焦式模仿学习。同时,恐惧信号被整合进策略优化与评估,作为安全性约束以降低适应过程中的性能退化。我们在模拟路线迁移任务、基于 Waymo 的回放驾驶场景、CARLA 障碍遮蔽危险以及真实移动机器人导航测试中评估 SAGE。在这些场景中,SAGE 提升了对新场景和高危情境的鲁棒性,减少了安全违规,且任务性能与强基线策略相当。结果表明,智能体可通过识别自身能力边界、适时请求指导、并从稀有高价值事件中选择性学习,在初始训练后实现持续改进。

原文摘要 · Abstract (English)

Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is insufficient, seek timely assistance, and convert safety-critical encounters into targeted improvement. Here we present self-aware guided exploration (SAGE), an active learning framework for post-training adaptation in AD. SAGE learns a predictive world model that generates two online intrinsic signals: fear, which estimates short-horizon predictive risk and model uncertainty, and curiosity, which measures novelty through prediction error. Curiosity adaptively calibrates the intervention threshold for fear, allowing the agent to regulate risk in a context-dependent manner. When predicted fear exceeds this adaptive threshold, the agent transfers control to an expert or fallback policy and uses the resulting takeover trajectories for focused imitation learning. In parallel, fear is integrated into policy optimization and evaluation as a safety-oriented constraint to reduce performance regressions during adaptation. We evaluate SAGE in simulated route-transfer tasks, Waymo-based logged driving scenarios, CARLA occlusion hazards, and real-world mobile robot navigation tests. Across these settings, SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies. These results suggest that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.

自动驾驶主动学习自感知持续进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。