arXiv:2412.02951cs.ROcs.LG2024-12

用强化学习让感知模型理解系统安全,提升自动驾驶可靠性。

Incorporating System-level Safety Requirements in Perception Models via Reinforcement Learning

  • 将安全规则转化为评分,融入强化学习奖励函数
  • 仿真显示新方法显著提升系统级安全性
  • 适合关注自动驾驶安全的开发者与研究者

自主系统中的感知模块通常独立于下游决策与控制模块进行开发和优化,依赖准确率、精确率、召回率等传统性能指标。传统的损失函数(如交叉熵、负对数似然)仅关注减少误分类错误,却未考虑这些错误对系统级安全的影响,忽视了不同错误导致的系统失败严重程度差异。为解决此问题,我们提出一种新型训练范式,通过将基于规则书形式化表达的系统级安全要求转化为安全评分,并将其纳入强化学习框架的奖励函数中,对感知模型进行微调以实现系统级安全目标。仿真结果表明,采用该方法训练的模型在系统级安全性能上优于基准感知模型。

原文摘要 · Abstract (English)

Perception components in autonomous systems are often developed and optimized independently of downstream decision-making and control components, relying on established performance metrics like accuracy, precision, and recall. Traditional loss functions, such as cross-entropy loss and negative log-likelihood, focus on reducing misclassification errors but fail to consider their impact on system-level safety, overlooking the varying severities of system-level failures caused by these errors. To address this limitation, we propose a novel training paradigm that augments the perception component with an understanding of system-level safety objectives. Central to our approach is the translation of system-level safety requirements, formally specified using the rulebook formalism, into safety scores. These scores are then incorporated into the reward function of a reinforcement learning framework for fine-tuning perception models with system-level safety objectives. Simulation results demonstrate that models trained with this approach outperform baseline perception models in terms of system-level safety.

感知模型强化学习自动驾驶系统安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。