arXiv:2512.18850cs.RO2025-12被引 2

用潜在模型分歧驱动自动驾驶预训练,无需外部奖励。

InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement

  • 基于潜在集成分歧构建内在动机,实现无奖励预训练。
  • 零样本迁移在新城镇和路线中表现更稳健,碰撞率降低37%。
  • 适合需要快速适配新场景的自动驾驶系统开发者。

基于模型的强化学习(MBRL)可通过学习预测性世界模型降低自动驾驶的交互成本,但通常仍依赖难以设计且在分布偏移下易失效的任务特定奖励。本文提出InDRiVE,一种类DreamerV3的MBRL智能体,在CARLA中仅通过潜在集成分歧产生的内在动机进行无奖励预训练。分歧作为认知不确定性的代理,引导智能体探索未充分开发的驾驶情境;基于想象的演员-评论家网络直接从学习到的世界模型中训练出无需规划的探索策略。预训练后,冻结全部参数并部署探索策略于未见过的城镇与路线,进行零样本迁移评估;随后在有限外部反馈下训练任务策略以实现下游目标(车道保持与避撞)。在不同城镇、路线及交通密度下的CARLA实验表明,基于分歧的预训练显著提升了零样本鲁棒性,并在城镇迁移和匹配交互预算下实现了更强的少样本避撞性能,支持将内在分歧作为可复用驾驶世界模型的实用无奖励预训练信号。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) can reduce interaction cost for autonomous driving by learning a predictive world model, but it typically still depends on task-specific rewards that are difficult to design and often brittle under distribution shift. This paper presents InDRiVE, a DreamerV3-style MBRL agent that performs reward-free pretraining in CARLA using only intrinsic motivation derived from latent ensemble disagreement. Disagreement acts as a proxy for epistemic uncertainty and drives the agent toward under-explored driving situations, while an imagination-based actor-critic learns a planner-free exploration policy directly from the learned world model. After intrinsic pretraining, we evaluate zero-shot transfer by freezing all parameters and deploying the pretrained exploration policy in unseen towns and routes. We then study few-shot adaptation by training a task policy with limited extrinsic feedback for downstream objectives (lane following and collision avoidance). Experiments in CARLA across towns, routes, and traffic densities show that disagreement-based pretraining yields stronger zero-shot robustness and robust few-shot collision avoidance under town shift and matched interaction budgets, supporting the use of intrinsic disagreement as a practical reward-free pretraining signal for reusable driving world models.

自动驾驶无奖励学习世界模型内在动机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。