让驾驶规划同时具备密集奖励监督和动态生成能力。
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

- 用流匹配模型学习奖励条件下的动作分布,统一两种规划方法优势。
- 在NAVSIM v1/v2上实现当前最优多模态规划结果,提案质量显著提升。
- 支持按需采样,适合需要安全与灵活性兼顾的自动驾驶场景。
多模态驾驶规划长期面临两种范式的矛盾:评分式方法依赖密集奖励监督,但受限于固定动作词表;锚点式方法可动态生成候选路径,却仅能获得单条真实轨迹的稀疏监督。本文提出FlowR2A,将基于仿真的奖励从判别目标重构为生成条件,通过流匹配解码器从密集的轨迹-奖励对中学习奖励条件下的动作分布,使模型在单一生成框架内兼具评分法的密集监督优势与锚点法的动态生成能力,从而内化动作与其在安全性、前进效率、舒适性及规则遵守方面的关联。为平衡硬性安全约束与软性前进目标,引入细粒度的时间步奖励条件与奖励噪声增强。生成式架构自然支持测试时通过奖励引导与锚点采样进行可控生成,产出高质量多模态提案。FlowR2A在NAVSIM v1和v2基准上取得当前最优性能,其多模态提案质量显著优于先前方法。
原文摘要 · Abstract (English)
Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。