提出新型3D占据预测方法,突破传统视角投影限制,提升全局理解与鲁棒性。
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

- 用分层张量压缩实现密集2D到3D的自由路由,打破局部性瓶颈。
- 在Occ3D-nuScenes和Occ3D-Waymo上达到当前最佳性能。
- 无需相机外参也能保持强结构鲁棒性,适合多视角无标定场景。
基于视觉的3D占据预测依赖于2D到3D的视图转换。现有方法主要采用显式物理投影,将路由矩阵严格限制在稀疏的相机光线上。虽计算高效,但造成严重局部性瓶颈,使网络难以构建全局上下文理解,且在相机外参不可靠或缺失时性能显著下降。为突破此瓶颈,我们抽象视图转换为无约束的二部图路由,提出因子化密集路由(FDR)。通过层次化张量收缩近似密集2D到3D混合,FDR保证全局限收场,同时复杂度可控且低于二次方。关键的是,密集路由中的必有空间压缩揭示了分辨率与上下文间的根本权衡。为此,我们设计分辨率-上下文解耦架构:将3D空间分解为全局宏观拓扑锚点(由FDR实现)与精确局部几何平面(由显式投影实现)。该解耦使全局语义推断与精确表面定位可互补共存而不相互牺牲。大量实验表明,本框架在Occ3D-nuScenes与Occ3D-Waymo基准上达到领先性能。尤为突出的是,在未标定设置下(物理外参缺失),其全局路由内化隐式多相机刚体拓扑,相较物理投影基线表现出更强的结构鲁棒性。
原文摘要 · Abstract (English)
Vision-based 3D occupancy prediction fundamentally relies on the 2D-to-3D view transformation. Current paradigms predominantly utilize explicit physical projection, which artificially restricts the routing matrix to strict, sparse camera rays. While computationally efficient, this imposes a severe Locality Bottleneck, preventing the network from constructing holistic contextual understanding and degrading sharply when camera extrinsics are unreliable or absent. To break this bottleneck, we abstract view transformation as unconstrained bipartite routing and propose Factorized Dense Routing (FDR). By approximating dense 2D-to-3D mixing through hierarchical tensor contractions, FDR guarantees a fully-global receptive field with tractable, sub-quadratic complexity. Crucially, the mandatory spatial contraction in dense routing exposes a fundamental Resolution-Context Trade-off. To address this, we introduce a Resolution-Context Decoupled Architecture. We factorize the 3D space into a global macroscopic topological anchor (via FDR) and precise local geometric planes (via explicit projection). This decoupling enables global semantic inference and exact surface localization to complement each other without mutual compromise. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on the Occ3D-nuScenes and Occ3D-Waymo benchmarks. More notably, in an uncalibrated setting where physical extrinsics are withheld, our global routing internalizes the implicit multi-camera rig topology and exhibits substantially stronger structural robustness than physical-projection baselines under the same protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。