用强化学习提升交通模拟真实度,还能精准控制场景生成。
RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-Tuning

- 基于强化学习微调预训练模型,让仿真更贴近真实数据分布。
- 在Waymo数据集上实现顶尖真实度,样本需求比传统方法少得多。
- 支持目标条件控制,适合需要可定制交通场景的研究与应用。
监督式开环训练广泛用于交通模拟模型训练,但难以捕捉复杂驾驶场景中固有的动态多智能体交互。本文提出RLFTSim,一种基于强化学习的微调框架,通过将仿真轨迹对齐真实世界数据分布来提升场景真实性,并实现目标条件化的可控性蒸馏。我们在预训练仿真模型基础上构建该框架,设计兼顾保真度与可控性的奖励函数,并在Waymo Open Motion Dataset上进行全面实验。结果表明,该方法在真实性上达到当前最优水平;相比其他启发式搜索微调方法,得益于低方差、密集的奖励信号,所需样本量显著减少,且从设计上直接解决真实性对齐问题。我们还验证了该方法在交通场景可控性蒸馏上的有效性。
原文摘要 · Abstract (English)
Supervised open-loop training has been widely adopted for training traffic simulation models; however, it fails to capture the inherently dynamic, multi-agent interactions common in complex driving scenarios. We introduce RLFTSim, a reinforcement-learning-based fine-tuning framework that enhances scenario realism by aligning simulator rollouts with real-world data distributions and provides a method for distilling goal-conditioned controllability in scenario generation. We instantiate RLFTSim on top of a pre-trained simulation model, design a reward that balances fidelity and controllability, and perform comprehensive experiments on the Waymo Open Motion Dataset. Our results show improvements in realism, achieving state-of-the-art performance. Compared with other heuristic search-based fine-tuning methods, RLFTSim requires significantly fewer samples due to a proposed low-variance and dense reward signal, and it directly addresses the realism alignment issue by design. We also demonstrate the effectiveness of our approach for distilling traffic simulation controllability through goal conditioning. The project page is available at https://ehsan-ami.github.io/rlftsim.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。