arXiv:2603.14076cs.CV2026-03被引 1

解决单目3D占位预测中深度模糊与冷启动问题,提升场景结构清晰度。

SGR-OCC: Evolving Monocular Priors for Embodied 3D Occupancy Prediction via Soft-Gating Lifting and Semantic-Adaptive Geometric Refinement

  • 用软门控机制抑制背景噪声,显式建模深度不确定性。
  • 局部预测完成率58.55%,语义分割准确率49.89%,优于前人方法。
  • 适合需要高精度3D场景理解的机器人导航任务。

3D语义占位预测是具身AI的核心,使智能体能从单目视频流中增量感知密集场景几何与语义。然而现有在线框架存在两大瓶颈:单目估计固有的深度模糊导致物体边界出现“特征渗漏”,以及初始化阶段时间融合层不稳定引发的“冷启动”问题,破坏高质量空间先验。本文提出SGR-OCC(Soft-Gating and Ray-refinement Occupancy),基于“继承与演化”理念构建统一框架。为完美继承单目空间能力,引入软门控特征提升模块,通过高斯门控概率性抑制背景噪声;同时设计动态射线约束锚点精修模块,将复杂3D偏移搜索简化为沿相机射线的一维深度修正,确保亚体素级贴合物理表面。为保障稳定演化以达成时间一致性,采用两阶段渐进训练策略,结合身份初始化融合机制,有效解决冷启动问题,保护空间先验免受早期噪声梯度干扰。在EmbodiedOcc-ScanNet和Occ-ScanNet基准上的大量实验表明,SGR-OCC达到顶尖性能:局部预测任务中完成率IoU达58.55%,语义mIoU为49.89%,较前最优方法EmbodiedOcc++分别提升3.65%和3.69%;在更具挑战的具身预测任务中,达到55.72% SC-IoU与46.22% mIoU。定性结果进一步验证了模型在复杂室内环境中保持结构完整性和边界锐度的卓越能力。

原文摘要 · Abstract (English)

3D semantic occupancy prediction is a cornerstone for embodied AI, enabling agents to perceive dense scene geometry and semantics incrementally from monocular video streams. However, current online frameworks face two critical bottlenecks: the inherent depth ambiguity of monocular estimation that causes "feature bleeding" at object boundaries , and the "cold start" instability where uninitialized temporal fusion layers distort high-quality spatial priors during early training stages. In this paper, we propose SGR-OCC (Soft-Gating and Ray-refinement Occupancy), a unified framework driven by the philosophy of "Inheritance and Evolution". To perfectly inherit monocular spatial expertise, we introduce a Soft-Gating Feature Lifter that explicitly models depth uncertainty via a Gaussian gate to probabilistically suppress background noise. Furthermore, a Dynamic Ray-Constrained Anchor Refinement module simplifies complex 3D displacement searches into efficient 1D depth corrections along camera rays, ensuring sub-voxel adherence to physical surfaces. To ensure stable evolution toward temporal consistency, we employ a Two-Phase Progressive Training Strategy equipped with identity-initialized fusion, effectively resolving the cold start problem and shielding spatial priors from noisy early gradients. Extensive experiments on the EmbodiedOcc-ScanNet and Occ-ScanNet benchmarks demonstrate that SGR-OCC achieves state-of-the-art performance. In local prediction tasks, SGR-OCC achieves a completion IoU of 58.55$\%$ and a semantic mIoU of 49.89$\%$, surpassing the previous best method, EmbodiedOcc++, by 3.65$\%$ and 3.69$\%$ respectively. In challenging embodied prediction tasks, our model reaches 55.72$\%$ SC-IoU and 46.22$\%$ mIoU. Qualitative results further confirm our model's superior capability in preserving structural integrity and boundary sharpness in complex indoor environments.

3D占位单目视觉具身智能几何精修

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。