用强化学习微调让交通模拟更像真人驾驶。
Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning
- 用奖励驱动方式优化模型行为,对齐人类驾驶偏好。
- 在Waymo数据集上实现0.7858的仿真真实度评分,领先榜单。
- 迭代式训练策略提升模型泛化能力,适合自动驾驶研究者。
大规模多智能体交通行为的可扩展且真实的模拟对推动自动驾驶技术至关重要。尽管现有数据驱动的模拟器已在该领域取得显著进展,但主要依赖监督学习来对齐模拟分布与真实驾驶场景。然而,训练与测试之间的分布偏移问题仍持续存在,常导致模型在未见环境中的泛化能力下降。为此,我们提出SMART-R1,一种专为下一词预测模型设计的R1风格强化微调范式,以更好地对齐智能体行为与人类偏好及评估指标。该方法引入面向指标的策略优化算法,提升分布对齐效果,并采用交替进行监督微调(SFT)与强化微调(RFT)的迭代训练策略,最大化性能提升。在大规模Waymo开放运动数据集(WOMD)上的大量实验验证了该简单而强大的R1风格训练框架在增强基础模型方面的有效性。在Waymo开放模拟智能体挑战赛(WOSAC)中,SMART-R1以总体真实度元分数0.7858取得领先成绩,提交时位居榜首。
原文摘要 · Abstract (English)
Scalable and realistic simulation of multi-agent traffic behavior is critical for advancing autonomous driving technologies. Although existing data-driven simulators have made significant strides in this domain, they predominantly rely on supervised learning to align simulated distributions with real-world driving scenarios. A persistent challenge, however, lies in the distributional shift that arises between training and testing, which often undermines model generalization in unseen environments. To address this limitation, we propose SMART-R1, a novel R1-style reinforcement fine-tuning paradigm tailored for next-token prediction models to better align agent behavior with human preferences and evaluation metrics. Our approach introduces a metric-oriented policy optimization algorithm to improve distribution alignment and an iterative "SFT-RFT-SFT" training strategy that alternates between Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT) to maximize performance gains. Extensive experiments on the large-scale Waymo Open Motion Dataset (WOMD) validate the effectiveness of this simple yet powerful R1-style training framework in enhancing foundation models. The results on the Waymo Open Sim Agents Challenge (WOSAC) showcase that SMART-R1 achieves state-of-the-art performance with an overall realism meta score of 0.7858, ranking first on the leaderboard at the time of submission.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。