arXiv:2510.11501cs.LGcs.RO2025-10中稿 · IEEE ICAR 2025被引 1

提出新方法提升自动驾驶赛车在未知对手下的泛化能力。

Context-Aware Model-Based Reinforcement Learning for Autonomous Racing

  • 用上下文建模对手行为,让算法适应不同驾驶风格。
  • 新方法cMask在未知对手下表现更好,对已知对手也更优。
  • 适合研究自动驾驶泛化与强化学习的开发者参考。

自动驾驶车辆有望显著提升道路安全。为实现真实应用,需具备良好泛化能力的算法。模型基于强化学习(MBRL)在多个领域表现出色且数据效率高,但对环境变化敏感。本文在模拟赛车环境Roboracer中研究MBRL在自主驾驶中的性能与泛化能力,将竞速任务建模为上下文马尔可夫决策过程,通过上下文参数化对手行为及环境动态。实验对比了多种MBRL算法,并提出新方法cMask。结果表明,上下文感知的MBRL在分布外对手行为下泛化能力更强;cMask在分布内对手上也表现更优,兼具强泛化性与更高性能。

原文摘要 · Abstract (English)

Autonomous vehicles have shown promising potential to be a groundbreaking technology for improving the safety of road users. For these vehicles, as well as many other safety-critical robotic technologies, to be deployed in real-world applications, we require algorithms that can generalize well to unseen scenarios and data. Model-based reinforcement learning algorithms (MBRL) have demonstrated state-of-the-art performance and data efficiency across a diverse set of domains. However, these algorithms have also shown susceptibility to changes in the environment and its transition dynamics. In this work, we explore the performance and generalization capabilities of MBRL algorithms for autonomous driving, specifically in the simulated autonomous racing environment, Roboracer (formerly F1Tenth). We frame the head-to-head racing task as a learning problem using contextual Markov decision processes and parameterize the driving behavior of the adversaries using the context of the episode, thereby also parameterizing the transition and reward dynamics. We benchmark the behavior of MBRL algorithms in this environment and propose a novel context-aware extension of the existing literature, cMask. We demonstrate that context-aware MBRL algorithms generalize better to out-of-distribution adversary behaviors relative to context-free approaches. We also demonstrate that cMask displays strong generalization capabilities, as well as further performance improvement relative to other context-aware MBRL approaches when racing against adversaries with in-distribution behaviors.

强化学习自动驾驶泛化能力模型预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。