arXiv:2605.19524cs.ROcs.CV2026-05被引 1

通过负样本提升自动驾驶模型对危险行为的理解能力

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

论文配图:SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving
图 1 · 摘自论文原文
  • 引入反事实推理生成风险场景的负样本和对应正轨迹
  • 在NAVSIM上碰撞率降至3.36%,语言理解准确率达84.2%
  • 适合关注安全边界建模与高风险行为抑制的研究者

端到端自动驾驶系统在常规场景表现优异,但在安全关键的长尾场景中仍存不足。视觉-语言-动作(VLA)模型因具备强推理能力而前景广阔,但多数方法仅依赖正向专家示范,忽视负样本,导致对危险行为和安全边界的理解不足。为此,本文提出SafeAlign-VLA,一种统一的负样本增强安全对齐框架,将负样本融入监督学习与强化学习。首先,设计反事实安全配对范式,通过反事实推理从风险场景生成结构化安全标签及反事实正向轨迹;其次采用两阶段训练:第一阶段为负样本增强的监督微调,用于失败反馈与轨迹修正;第二阶段为基于锚点的组相对策略优化,以正负轨迹作为对比锚点,通过组相对优势引导采样并惩罚高风险行为。在NAVSIM与DeepAccident数据集上的实验验证了该框架的有效性。SafeAlign-VLA在NAVSIM v1测试集上达到89.1 PDMS,较无负样本基线提升1.3%;在DeepAccident上碰撞率降至3.36%,同时实现84.2%的语言准确率与85.8%的风险预测准确率。

原文摘要 · Abstract (English)

End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models are promising due to their strong reasoning capabilities. However, most VLA-based approaches rely on positive expert demonstrations, rarely exploiting negative samples, leading to insufficient understanding of risky behaviors and safety boundaries. To address this limitation, we propose SafeAlign-VLA, a unified negative-enhanced safe alignment framework that incorporates negative data into supervised learning and reinforcement learning. First, we develop a counterfactual safety pairing paradigm to generate structured safety labels and counterfactual positive trajectories from risky scenarios via counterfactual reasoning. Then, a two-stage training strategy is adopted: negative-enhanced supervised fine-tuning for failure feedback and trajectory correction, followed by anchor-based group relative policy optimization that uses positive and negative trajectories as contrastive anchors to steer sampling and penalize high-risk behaviors via group-relative advantages. Experiments on NAVSIM and DeepAccident validate the proposed framework. SafeAlign-VLA achieves 89.1 PDMS on the NAVSIM v1 testset, improving over the baseline without negative data by 1.3%. On DeepAccident, it reduces the collision rate to 3.36%, while achieving 84.2% language accuracy and 85.8% risk prediction accuracy. These results demonstrate the effectiveness of the proposed negative-enhanced safe alignment framework for safe and robust autonomous driving.

自动驾驶安全对齐负样本VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。