用一张全景图实现开放词汇的3D占位预测,让机器人在未知环境中看得更全更准。
O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Embodied Intelligent Robotics
- 通过极坐标螺旋结构建模360°全景体素,支持连续空间表示与长程上下文捕捉。
- 在QuadOcc和Human360Occ上达到顶尖性能,且跨场景泛化能力强。
- 无需梯度优化的视觉-体素-文本对齐机制,适合开放世界机器人感知任务。
消费电子向具身智能演进加速了消费级具身智能机器人(CEIRs)的发展,这类设备需在复杂真实环境中感知、理解并交互。通过全方位感知重建3D世界对在复杂动态环境中运行的CEIRs日益重要。然而,现有基于视觉的3D占位预测方法受限于有限视角输入和预定义训练分布,难以支持需全面安全感知的开放世界探索。为此,我们提出O3N,首个从单张全景RGB图像进行开放词汇占位预测的框架。O3N通过极坐标螺旋Mamba(PsM)模块将全景体素嵌入极坐标螺旋拓扑,实现360°连续空间表示与长程上下文建模。占位成本聚合(OCA)模块在体素空间内统一几何与语义监督,确保重构几何与底层语义结构一致。自然模态对齐(NMA)建立无梯度对齐路径,调和视觉特征、体素嵌入与文本语义,形成一致的像素-体素-文本表示三元组,支持开放世界感知。多模型实验证明,本方法在QuadOcc和Human360Occ基准上达到当前最优表现,并展现显著的跨场景泛化与语义可扩展性。源代码将公开于https://github.com/MengfeiD/O3N。
原文摘要 · Abstract (English)
The rapid evolution of consumer electronics toward embodied intelligence has accelerated the emergence of Consumer Embodied Intelligent Robotics (CEIRs), where intelligent devices are expected to perceive, understand, and interact with complex real-world environments. Understanding and reconstructing the 3D world through omnidirectional perception is therefore becoming increasingly important for CEIRs operating in complex and dynamic environments. However, existing vision-based 3D occupancy prediction methods are constrained by limited perspective inputs and a predefined training distribution, making them difficult to support embodied intelligent systems that require comprehensive and safe perception of scenes in open-world exploration. To address this, we present O3N, the first framework for open-vocabulary occupancy prediction from a single omnidirectional RGB image. O3N embeds omnidirectional voxels in a polar-spiral topology via the Polar-spiral Mamba (PsM) module, enabling continuous spatial representation and long-range context modeling across 360°. The Occupancy Cost Aggregation (OCA) module introduces a principled mechanism for unifying geometric and semantic supervision within the voxel space, ensuring consistency between reconstructed geometry and underlying semantic structure. Moreover, Natural Modality Alignment (NMA) establishes a gradient-free alignment pathway that harmonizes visual features, voxel embeddings, and text semantics, forming a consistent pixel-voxel-text representation triad for open-world perception. Extensive experiments on multiple models demonstrate that our method not only achieves state-of-the-art performance on QuadOcc and Human360Occ benchmarks but also exhibits remarkable cross-scene generalization and semantic scalability. The source code will be made publicly available at https://github.com/MengfeiD/O3N.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。