让人类实时指导自动驾驶模型,提升复杂路况下的决策能力。
Learning from Human Driving: A Human-in-the-Loop Online Behavior Cloning Framework for Autonomous Driving

- 通过人类干预在线优化驾驶策略,融合大模型感知与人类智能
- 在CARLA测试中驾驶表现提升超47%,极端场景更安全可靠
- 适合研究自动驾驶交互决策与人机协同的团队参考
随着大基础模型(LFMs)的发展,数据驱动的自动驾驶取得了显著进展。然而,现有方法在复杂交互和长尾场景中仍面临分布偏移与因果混淆问题,导致决策灵活性与极端条件下的安全性不足。为此,本文提出一种人机协同在线行为克隆框架(HiL-OBC),旨在深度融合大模型的跨模态感知能力与人类专家的高层驾驶智慧。该框架包含三个关键阶段:人工干预下的策略初始化、基于贝叶斯策略适应的潜在行为建模,以及在线部署与更新。此外,设计了多模态在线行为克隆(MOBC)模型,通过轻量网络结构、接管触发机制和多变损失函数,实现对基础驾驶策略的在线优化,增强复杂环境中的决策鲁棒性。在LangAuto-Human CARLA基准上评估表明,经人机协同机制优化的驾驶策略性能显著提升:StructNav、LFG、LMDrive的DS分别提高47.25%、31.59%、32.12%,且在多种实验设置与组件分析中验证了人机协同学习在提升决策鲁棒性与整体驾驶性能方面的优势。
原文摘要 · Abstract (English)
With the evolution of large foundation models (LFMs), data-driven autonomous driving has made significant strides. However, existing paradigms still face severe challenges in complex interaction and long-tail scenarios due to distribution shift and causal confusion. These limitations often result in a lack of human-level decision-making flexibility and safety in extreme conditions. To overcome this limitation, this paper proposes a Human-in-the-Loop Online Behavior Cloning frame work (HiL-OBC) for autonomous driving, which aims to deeply integrate the cross-modal perceptual capabilities of LFMs with the high-level driving intelligence of human experts. Specifically, HiL-OBC deployment is executed through three critical phases: policy initialization with human intervention, latent behavioral modeling with Bayesian policy adaptation, and online deploy ment and updates. Furthermore, we design a Multi-modal Online Behavior Cloning (MOBC) model, which optimizes the base driving policy online through a lightweight network architecture, a takeover trigger mechanism, and a multi-variant loss function, thereby enhancing the system's decision-making robustness in complex environments. We evaluated the HiL-OBC on the LangAuto-Human CARLA benchmark. Experimental results demonstrate that the driving policies optimized via the human-in-the-loop mechanism achieve substantial performance gains: the DS of StructNav, LFG, and LMDrive increased by 47.25%, 31.59%, and 32.12%, respectively, with a simultaneous of various experimental settings and key components highlights the advantages of human-in-the-loop learning in improving decision-making robustness and overall driving performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。