arXiv:2512.13262cs.ROcs.CV2025-12被引 2

通过后训练与测试时扩展,让自动驾驶模型更安全、更反应灵敏。

Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving

  • 用人类监督的强化学习方法微调预训练模型,提升安全性。
  • 仅用10%数据,安全性能提升超40%,轨迹仍自然真实。
  • 测试时动态调整行为选择,适合实际驾驶场景部署。

学习多智能体间的交互运动行为是自动驾驶的核心挑战。虽然模仿学习模型能生成逼真轨迹,但常继承以安全示范为主的训练数据偏差,在关键安全场景下鲁棒性不足。此外,多数研究依赖开环评估,忽略了闭环执行中的误差累积。为此,我们提出两种互补策略:首先,提出组相对行为优化(GRBO),一种基于群体相对优势最大化与人类正则化的强化学习后训练方法。仅使用10%训练数据,GRBO使安全性能提升超过40%,同时保持行为真实性。其次,引入温启动Top-K采样(Warm-K),在测试时平衡行为一致性与多样性。该方法无需重训练即可提升测试时行为的一致性与响应速度,缓解协变量偏移,降低性能差异。演示视频见补充材料。

原文摘要 · Abstract (English)

Learning interactive motion behaviors among multiple agents is a core challenge in autonomous driving. While imitation learning models generate realistic trajectories, they often inherit biases from datasets dominated by safe demonstrations, limiting robustness in safety-critical cases. Moreover, most studies rely on open-loop evaluation, overlooking compounding errors in closed-loop execution. We address these limitations with two complementary strategies. First, we propose Group Relative Behavior Optimization (GRBO), a reinforcement learning post-training method that fine-tunes pretrained behavior models via group relative advantage maximization with human regularization. Using only 10% of the training dataset, GRBO improves safety performance by over 40% while preserving behavioral realism. Second, we introduce Warm-K, a warm-started Top-K sampling strategy that balances consistency and diversity in motion selection. Our Warm-K method-based test-time scaling enhances behavioral consistency and reactivity at test time without retraining, mitigating covariate shift and reducing performance discrepancies. Demo videos are available in the supplementary material.

自动驾驶行为建模强化学习测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。