构建首个面向交互式驾驶世界模型的多样化数据集
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
- 专为训练交互式驾驶世界模型设计,包含完整驾驶操作
- 涵盖多智能体复杂互动与开放世界知识,视频多样性显著提升
- 提供动作指令跟随评测基准,适合自动驾驶研究者使用
驾驶世界模型因其建模复杂物理动态的能力而受到越来越多关注。然而,当前驾驶数据集视频多样性有限,制约了其建模能力的充分发挥。我们提出 DrivingDojo,首个专为训练交互式世界模型设计的驾驶数据集,包含完整的驾驶操作、多样化的多智能体互动以及丰富的开放世界驾驶知识,为未来世界模型发展奠定基础。我们进一步定义了一个动作指令跟随(AIF)评测基准,验证了该数据集在生成动作可控未来预测方面的优越性。
原文摘要 · Abstract (English)
Driving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashed due to the limited video diversity in current driving datasets. We introduce DrivingDojo, the first dataset tailor-made for training interactive world models with complex driving dynamics. Our dataset features video clips with a complete set of driving maneuvers, diverse multi-agent interplay, and rich open-world driving knowledge, laying a stepping stone for future world model development. We further define an action instruction following (AIF) benchmark for world models and demonstrate the superiority of the proposed dataset for generating action-controlled future predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。