首个可处理未标定城市场景的3D占位预测模型,支持多视角输入。
OccAny: Generalized Unconstrained Urban 3D Occupancy
- 基于视觉几何基础模型,实现无约束城市场景的3D占位预测
- 引入分割强制机制,提升占位质量并支持像素级分割输出
- 设计新视角渲染流程,通过测试时视图增强完成复杂场景几何补全
现有3D占位预测方法依赖域内标注和精确传感器先验,难以扩展且泛化能力差。尽管近期视觉几何基础模型具备强泛化能力,但主要面向通用任务,缺乏城市占位预测所需的关键要素:度量预测、杂乱场景下的几何补全及城市场景适配性。本文提出OccAny,首个可在域外未标定场景中运行的无约束城市3D占位模型,能同时预测度量占位与分割特征。该模型灵活支持序列、单目或环视图像输入。贡献包括:(i) 首个通用3D占位框架;(ii) 分割强制机制,在提升占位质量的同时支持掩码级预测;(iii) 新视角渲染流水线,通过测试时视图增强实现几何补全。大量实验表明,OccAny在三个输入设置下均超越所有视觉几何基线,在两个主流城市占位数据集上性能媲美域内自监督方法。代码已开源。
原文摘要 · Abstract (English)
Relying on in-domain annotations and precise sensor-rig priors, existing 3D occupancy prediction methods are limited in both scalability and out-of-domain generalization. While recent visual geometry foundation models exhibit strong generalization capabilities, they were mainly designed for general purposes and lack one or more key ingredients required for urban occupancy prediction, namely metric prediction, geometry completion in cluttered scenes and adaptation to urban scenarios. We address this gap and present OccAny, the first unconstrained urban 3D occupancy model capable of operating on out-of-domain uncalibrated scenes to predict and complete metric occupancy coupled with segmentation features. OccAny is versatile and can predict occupancy from sequential, monocular, or surround-view images. Our contributions are three-fold: (i) we propose the first generalized 3D occupancy framework with (ii) Segmentation Forcing that improves occupancy quality while enabling mask-level prediction, and (iii) a Novel View Rendering pipeline that infers novel-view geometry to enable test-time view augmentation for geometry completion. Extensive experiments demonstrate that OccAny outperforms all visual geometry baselines on 3D occupancy prediction task, while remaining competitive with in-domain self-supervised methods across three input settings on two established urban occupancy prediction datasets. Our code is available at https://github.com/valeoai/OccAny .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。