arXiv:2603.13108cs.ROcs.CV2026-03被引 2

为四足机器人设计首个全景多模态占位数据集与感知框架,提升复杂环境下的3D感知能力。

Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots

  • 提出VoxelHound框架,融合全景视觉与多模态信息,适配腿式机器人运动特性。
  • 引入垂直抖动补偿模块,有效缓解机体俯仰/滚转导致的视角扰动,提升空间推理一致性。
  • 在PanoMMOcc数据集上实现mIoU提升4.16,适合研究具身机器人3D感知的研究者。

全景图像为四足机器人提供了360°环境感知的全面视图。然而,现有占位预测方法主要针对轮式自动驾驶设计,严重依赖RGB信息,在复杂动态环境中鲁棒性不足。为此,我们构建了首个真实世界的全景多模态占位数据集PanoMMOcc,涵盖四种传感模态,覆盖多样场景。同时提出VoxelHound框架,专为腿式运动与球面成像设计。该框架包含垂直抖动补偿(VJC)模块,可缓解运动中因机体俯仰/滚转带来的严重视角扰动,实现更一致的空间推理;以及多模态信息提示融合(MIPF)模块,高效整合全景视觉与辅助模态,提升体素级占位预测性能。我们在PanoMMOcc上建立了全面基准,并提供详细数据分析,支持对具身感知挑战场景的系统评估。大量实验表明,VoxelHound在PanoMMOcc上达到当前最优表现,mIoU提升4.16。数据集与代码将开源,以促进未来具身机器人全景多模态3D感知研究。

原文摘要 · Abstract (English)

Panoramic imagery provides holistic 360° visual coverage for environmental perception in quadruped robots. However, existing occupancy prediction methods are primarily designed for wheeled autonomous driving and rely heavily on RGB cues, which limits their robustness in complex, dynamically changing environments. To bridge this gap, we introduce PanoMMOcc, the first real-world panoramic multimodal occupancy dataset for quadruped robots, comprising four sensing modalities collected across diverse scenes. We further propose VoxelHound, a panoramic multimodal occupancy perception framework tailored to legged locomotion and spherical imaging. VoxelHound incorporates a Vertical Jitter Compensation (VJC) module to mitigate severe viewpoint perturbations caused by body pitch and roll during locomotion, enabling more consistent spatial reasoning, and a Multimodal Information Prompt Fusion (MIPF) module to effectively integrate panoramic visual cues with auxiliary modalities for enhanced volumetric occupancy prediction. We also establish a comprehensive benchmark on PanoMMOcc and provide detailed dataset analyses to enable systematic evaluation in challenging embodied perception scenarios. Extensive experiments demonstrate that VoxelHound achieves state-of-the-art performance on PanoMMOcc, with a +4.16 gain in mIoU. The dataset and code will be publicly released to facilitate future research on panoramic multimodal 3D perception for embodied robotic systems at https://github.com/SXDR/PanoMMOcc.

四足机器人占位预测多模态感知全景视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。