arXiv:2503.06611cs.LGcs.SY2025-03被引 1

用逆强化学习生成自动驾驶最小暴露路径,提升数据合成能力。

Inverse Reinforcement Learning for Minimum-Exposure Paths in Spatiotemporally Varying Scalar Fields

  • 基于逆强化学习构建模型,从少量路径数据中学习最优避险策略。
  • 在相同或未知威胁场下均能生成低暴露路径,误差小且泛化能力强。
  • 适合用于自动驾驶安全评估与复杂环境行为数据增强。

自动驾驶车辆的性能与可靠性分析可借助工具将小规模数据集扩展为更多合理的行驶行为样本。本文聚焦于在固定目标位置下最小化车辆对不利环境条件的暴露问题。环境由一个正标量场表征,数值越高代表越危险的条件。本文提出一种逆强化学习(IRL)模型,用于生成与训练数据相似的最小暴露路径。该模型适用于静态与动态变化的威胁场。实验表明,当威胁场与训练时一致时,模型能有效生成训练数据外初始状态下的路径;在未见过的威胁场上也表现出低误差。此外,模型在不同特征的数据集上训练后,可生成具有差异性的新路径数据集。

原文摘要 · Abstract (English)

Performance and reliability analyses of autonomous vehicles (AVs) can benefit from tools that ``amplify'' small datasets to synthesize larger volumes of plausible samples of the AV's behavior. We consider a specific instance of this data synthesis problem that addresses minimizing the AV's exposure to adverse environmental conditions during travel to a fixed goal location. The environment is characterized by a threat field, which is a strictly positive scalar field with higher intensities corresponding to hazardous and unfavorable conditions for the AV. We address the problem of synthesizing datasets of minimum exposure paths that resemble a training dataset of such paths. The main contribution of this paper is an inverse reinforcement learning (IRL) model to solve this problem. We consider time-invariant (static) as well as time-varying (dynamic) threat fields. We find that the proposed IRL model provides excellent performance in synthesizing paths from initial conditions not seen in the training dataset, when the threat field is the same as that used for training. Furthermore, we evaluate model performance on unseen threat fields and find low error in that case as well. Finally, we demonstrate the model's ability to synthesize distinct datasets when trained on different datasets with distinct characteristics.

逆强化学习自动驾驶路径规划数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。