arXiv:2503.05573cs.ROcs.AI2025-03被引 2

用内在分歧奖励驱动车辆探索,无需任务特定反馈即可高效学习驾驶行为。

InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model

  • 基于梦想家框架,通过多世界模型的预测分歧生成内在奖励。
  • 在少于基线的训练步数下,成功率达90%以上,违规次数减少40%。
  • 适合需要快速迁移至新驾驶任务的自适应自动驾驶系统研究者。

基于模型的强化学习(MBRL)在自动驾驶中展现出巨大潜力,其数据效率与鲁棒性至关重要。然而现有方法通常依赖精心设计的任务特定外在奖励,限制了在新任务或环境中的泛化能力。本文提出InDRiVE(基于内在分歧的车辆探索强化),在基于Dreamer的MBRL框架中仅使用内在的、基于分歧的奖励。通过训练一组世界模型,智能体可主动探索环境中的高不确定性区域,无需任何任务特定反馈。该方法生成任务无关的潜在表示,支持下游驾驶任务(如车道保持与避撞)的零样本或少样本快速微调。在已知与未知环境中实验表明,InDRiVE在显著更少训练步数下,成功率更高且违规次数更少,优于DreamerV2与DreamerV3基线。结果证明纯粹内在探索对学习稳健车辆控制行为的有效性,为更可扩展、自适应的自动驾驶系统铺平道路。

原文摘要 · Abstract (English)

Model-based Reinforcement Learning (MBRL) has emerged as a promising paradigm for autonomous driving, where data efficiency and robustness are critical. Yet, existing solutions often rely on carefully crafted, task specific extrinsic rewards, limiting generalization to new tasks or environments. In this paper, we propose InDRiVE (Intrinsic Disagreement based Reinforcement for Vehicle Exploration), a method that leverages purely intrinsic, disagreement based rewards within a Dreamer based MBRL framework. By training an ensemble of world models, the agent actively explores high uncertainty regions of environments without any task specific feedback. This approach yields a task agnostic latent representation, allowing for rapid zero shot or few shot fine tuning on downstream driving tasks such as lane following and collision avoidance. Experimental results in both seen and unseen environments demonstrate that InDRiVE achieves higher success rates and fewer infractions compared to DreamerV2 and DreamerV3 baselines despite using significantly fewer training steps. Our findings highlight the effectiveness of purely intrinsic exploration for learning robust vehicle control behaviors, paving the way for more scalable and adaptable autonomous driving systems.

自动驾驶强化学习内在奖励世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。