arXiv:2507.13162cs.CVcs.AI2025-07NeurIPS被引 18

仅用469M参数和280小时视频,实现自动驾驶长时预测新标杆

Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models

  • 采用简洁设计,无需地图、深度图等额外传感器
  • 在转弯、城市交通等难题场景下表现领先,优于连续模型
  • 适合关注高效自动驾驶世界模型的开发者与研究者

现有自动驾驶世界模型在长时程生成与复杂场景泛化方面表现不佳。本文提出一种仅含469M参数、基于280小时视频训练的模型,不依赖地图、深度信息或多摄像头等额外监督信号。实验表明,该模型在转弯操作和城市交通等挑战性场景中性能达到当前最优。我们对比了离散令牌与连续流匹配模型的优劣,构建兼容两者的混合分词器进行对照实验。结果支持连续自回归模型:其对设计变化更鲁棒,且整体表现优于离散令牌模型。代码、模型及可视化结果已公开于https://lmb-freiburg.github.io/orbis.github.io/。

原文摘要 · Abstract (English)

Existing world models for autonomous driving struggle with long-horizon generation and generalization to challenging scenarios. In this work, we develop a model using simple design choices, and without additional supervision or sensors, such as maps, depth, or multiple cameras. We show that our model yields state-of-the-art performance, despite having only 469M parameters and being trained on 280h of video data. It particularly stands out in difficult scenarios like turning maneuvers and urban traffic. We test whether discrete token models possibly have advantages over continuous models based on flow matching. To this end, we set up a hybrid tokenizer that is compatible with both approaches and allows for a side-by-side comparison. Our study concludes in favor of the continuous autoregressive model, which is less brittle on individual design choices and more powerful than the model built on discrete tokens. Code, models and qualitative results are publicly available at https://lmb-freiburg.github.io/orbis.github.io/.

自动驾驶世界模型长时预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。