用熵引导采样提升交通模拟多样性与安全性
Learning Rollout from Sampling:An R1-Style Tokenized Traffic Simulation Model
- 基于运动标记熵设计自适应采样机制,挖掘高不确定性的潜在轨迹
- 通过群体相对策略优化实现多智能体行为的多样化与安全性提升
- 适合自动驾驶仿真、交通行为建模等需要高多样性的研究场景
从人类驾驶示范中学习多样且高保真的交通模拟对自动驾驶评估至关重要。当前主流的下一标记预测(NTP)范式虽通过监督微调实现迭代优化,但限制了对次优区域中潜在有价值运动标记的主动探索。我们提出新型标记化交通模拟策略R1Sim,首次尝试基于运动标记熵模式的强化学习,并系统分析不同运动标记对模拟结果的影响。具体而言,引入熵引导的自适应采样机制,聚焦此前被忽略的高不确定性但高潜力的运动标记;同时采用群组相对策略优化(GRPO),结合安全感知奖励设计优化运动行为。上述组件协同实现探索与利用的平衡,通过多样化的高不确定性采样和群组比较估计,生成真实、安全且多样的多智能体行为。在Waymo Sim Agent基准上的大量实验表明,R1Sim性能媲美现有最优方法。
原文摘要 · Abstract (English)
Learning diverse and high-fidelity traffic simulations from human driving demonstrations is crucial for autonomous driving evaluation. The recent next-token prediction (NTP) paradigm, widely adopted in large language models (LLMs), has been applied to traffic simulation and achieves iterative improvements via supervised fine-tuning (SFT). However, such methods limit active exploration of potentially valuable motion tokens, particularly in suboptimal regions. Entropy patterns provide a promising perspective for enabling exploration driven by motion token uncertainty. Motivated by this insight, we propose a novel tokenized traffic simulation policy, R1Sim, which represents an initial attempt to explore reinforcement learning based on motion token entropy patterns, and systematically analyzes the impact of different motion tokens on simulation outcomes. Specifically, we introduce an entropy-guided adaptive sampling mechanism that focuses on previously overlooked motion tokens with high uncertainty yet high potential. We further optimize motion behaviors using Group Relative Policy Optimization (GRPO), guided by a safety-aware reward design. Overall, these components enable a balanced exploration-exploitation trade-off through diverse high-uncertainty sampling and group-wise comparative estimation, resulting in realistic, safe, and diverse multi-agent behaviors. Extensive experiments on the Waymo Sim Agent benchmark demonstrate that R1Sim achieves competitive performance compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。