统一处理室内外3D占据,一个模型搞定不同场景和相机配置。
OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

- 以像素为中心的高斯框架,动态适配不同相机与场景尺度。
- 在Occ-ScanNet上达59.92% mIoU,SurroundOcc-nuScenes上达23.06% mIoU。
- 适合需要跨场景通用3D理解的机器人、自动驾驶应用。
3D占据预测是场景理解的基础,但现有方法通常针对特定场景类型和占据协议。本文提出跨场景3D语义占据预测新任务,要求单一模型同时处理异构的室内外场景,涵盖不同相机、空间范围、体素规格及语义分类体系。该设定带来核心挑战:在多变相机配置与场景尺度下实现度量一致且场景自适应的图像到3D映射。为此,我们提出OccAnyScene,基于预训练深度基础模型的像素-视锥中心高斯框架。具体通过像素对齐视锥特征聚合构建每个特征像素的相机感知视锥查询,并利用视锥参数化高斯生成将查询解码为多个受预测像素深度及对应视锥几何约束的位置与大小的高斯。该方法在室内Occ-ScanNet上取得59.92% mIoU,在室外SurroundOcc-nuScenes上达到23.06% mIoU,刷新当前最佳性能。
原文摘要 · Abstract (English)
3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。