arXiv:2606.27644cs.CV2026-06中稿 · IEEE Signal Proces…

用分层向量量化提升3D占位表示,让自动驾驶更懂场景结构。

CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

论文配图:CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations
图 1 · 摘自论文原文
  • 分步细化:从粗到细逐级重构3D场景结构。
  • 在4D占位预测与路径规划中超越纯视觉方法。
  • 适合关注场景本质表征的自动驾驶研究者。

本文提出CascadeOcc,一种新型占位世界模型,强调内在结构层次而非外部辅助模态在自动驾驶中的作用。占位世界模型通过预测未来驾驶环境并规划行驶轨迹,有效连接感知与决策,但现有方法常过度依赖外部模态或大语言模型,未能充分挖掘占位表示自身的内在潜力。为增强复杂3D场景的表征能力,我们引入分层向量量化(VQ)机制至自回归框架中,遵循由粗到细的原则,通过多尺度架构逐步细化细节。同时,结合TimeMixer捕捉多尺度时间依赖性,在空间与时间上建立双层次机制。在4D占位预测与运动规划基准测试中,CascadeOcc展现出优于现有视觉主导方法的性能,验证了优化内在表示是替代依赖外部基础模型的有效路径。

原文摘要 · Abstract (English)

This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for autonomous driving. Occupancy world models -- forecasting the future driving environment and planning the driving trajectory -- effectively bridge perception and planning, but current approaches often heavily rely on external modalities or large language models, failing to fully exploit the inherent structural potential of occupancy representations themselves. To enhance representational capacity for complex 3D scenes, we integrate a cascaded Vector Quantized (VQ) mechanism into an autoregressive framework. Following a coarse-to-fine principle, CascadeOcc progressively refines fine-grained details from global structures through a multi-scale architecture. Additionally, we incorporate a TimeMixer to capture multi-scale temporal dependencies, establishing a dual-hierarchy mechanism in both space and time. Experimental results on 4D occupancy forecasting and motion planning benchmarks demonstrate that CascadeOcc achieves superior performance among vision-centric approaches, validating that optimizing inherent representations is a powerful alternative to relying on external foundation models.

3D占位自动驾驶分层表示自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。