arXiv:2511.22039cs.CV2025-11中稿 · CVPR被引 5

用稀疏表示直接预测3D场景未来占据,性能超越现有方法。

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

  • 跳过图像到鸟瞰图的转换,直接从原始图像特征预测多帧占据
  • 在nuScenes数据集上1-3秒预测任务达到最新性能,显著领先
  • 支持任意未来轨迹条件,适合自动驾驶场景建模

本文提出一种新的轨迹条件化未来3D场景占据预测架构。与依赖变分自编码器(VAE)生成离散占据标记、限制表达能力的方法不同,本方法直接从原始图像特征端到端预测多帧未来占据。受GPT、VGGT等基础视觉与语言模型中注意力机制成功的启发,采用稀疏占据表示,避免了中间鸟瞰图(BEV)投影及其显式几何先验。该设计使Transformer更有效地捕捉时空依赖关系。通过规避离散标记化的有限容量和BEV表示的结构局限,本方法在nuScenes基准上1-3秒占据预测任务中达到当前最优性能,显著优于现有方法。此外,其展现出对场景动态的鲁棒理解,在任意未来轨迹条件下均保持高精度。

原文摘要 · Abstract (English)

This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs) to generate discrete occupancy tokens, which inherently limit representational capacity, our approach predicts multi-frame future occupancy in an end-to-end manner directly from raw image features. Inspired by the success of attention-based transformer architectures in foundational vision and language models such as GPT and VGGT, we employ a sparse occupancy representation that bypasses the intermediate bird's eye view (BEV) projection and its explicit geometric priors. This design allows the transformer to capture spatiotemporal dependencies more effectively. By avoiding both the finite-capacity constraint of discrete tokenization and the structural limitations of BEV representations, our method achieves state-of-the-art performance on the nuScenes benchmark for 1-3 second occupancy forecasting, outperforming existing approaches by a significant margin. Furthermore, it demonstrates robust scene dynamics understanding, consistently delivering high accuracy under arbitrary future trajectory conditioning.

占据预测3D建模轨迹条件Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。