arXiv:2608.30122cs.CVcs.AI2026-08

解决多轨迹监督与策略优化不匹配问题,提升自动驾驶安全性和成功率。

Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving

论文配图:Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving
图 1 · 摘自论文原文
  • 用帕累托最优筛选有效轨迹,过滤冲突样本
  • 在单轨迹推理下达成91.4和89.1的高分表现
  • 适合追求自动驾驶安全与鲁棒性的研究者

视觉-语言-动作(VLA)驾驶方法结合多轨迹模仿学习与组相对策略优化(GRPO),轨迹选择对性能至关重要。然而,部分高分轨迹虽提升模仿效果,却因优势估计偏离当前策略可行行为分布,导致后续优化偏离安全合规行为。为此,我们提出一种新框架,将多轨迹监督与策略优化对齐。通过将增强轨迹约束于真实可行区域邻域,并以帕累托最优替代传统综合评分,仅保留非支配候选,从源头过滤冲突样本。为确保扩展监督有效融入策略优化,引入两个互补机制:可行性优先优势分配,根据每组轨迹的可行性调整信用分配,引导完全不可行组趋向安全参考;动态蒸馏,跨优化轮次更新教师轨迹,持续传递有用监督。二者协同逐步将扩展监督收益转化为策略改进。在NAVSIM v1和v2上,单轨迹推理下分别取得91.4 PDMS和89.1 EPDMS,恢复658个初始失败场景中的440个,较原GRPO基线提升11.1%。

原文摘要 · Abstract (English)

Vision-language-action (VLA) driving methods increasingly combine multi-trajectory imitation learning with group-relative policy optimization (GRPO), making trajectory selection critical to final performance. However, some high-scoring trajectories that improve imitation can degrade subsequent GRPO by inducing advantage estimates misaligned with the current policy's feasible behavior distribution, driving updates away from safe and compliant behaviors. To address this, we propose a novel framework that aligns multi-trajectory supervision with policy optimization. To address the policy gradient bias induced by infeasible noisy trajectories outside the feasible region, augmented trajectories are constrained to a neighboring manifold of the ground-truth feasible region, and a Pareto-optimality criterion is adopted in place of the conventional aggregate score, retaining only non-dominated candidates and thereby filtering out conflicting samples at the source. To ensure that expanded trajectory supervision is effectively absorbed during policy optimization, we introduce two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation. The former adapts Pareto credit to the feasibility composition of each rollout group and guides fully infeasible groups toward safe references. The latter updates teacher trajectories across refinement rounds to continually transfer useful supervision. Together, they progressively translate the benefits of expanded supervision into policy improvement. On NAVSIM v1 and v2, our method achieves 91.4 PDMS and 89.1 EPDMS, respectively, under single-trajectory inference, and recovers 440 of 658 initially failed scenes, 11.1\% higher than the original GRPO baseline.

自动驾驶策略优化多轨迹学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。