arXiv:2511.03571cs.ROcs.CV2025-11中稿 · CVPR被引 8

用单个全景摄像头实现腿式机器人360°语义占位感知,解决行走抖动问题。

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

  • 融合环形全景与球面展开图,保持360°连续性与网格对齐
  • 在笛卡尔与柱坐标空间推理,减少离散化偏差,边界更清晰
  • 轻量级设计适配腿式机器人,支持无额外传感器的运动补偿

稳健的3D语义占位对于腿式/人形机器人至关重要,但多数语义场景补全(SSC)系统针对前向传感器的轮式平台。本文提出OneOcc,一种仅依赖视觉的全景式SSC框架,专为步态引起的体位抖动和360°连续性设计。OneOcc结合:(i) 双投影融合(DP-ER),利用环形全景及其等距展开图,保留360°连续性与网格对齐;(ii) 双网格体素化(BGV),在笛卡尔与柱面极坐标空间中推理,降低离散化偏差并锐化自由/占据边界;(iii) 轻量级解码器搭配分层AMoE-3D,实现动态多尺度融合,提升长距离与遮挡推理能力;(iv) 即插即用的步态位移补偿(GDC),在特征层面学习运动校正,无需额外传感器。我们还发布了两个全景占位基准:QuadOcc(真实四足机器人,第一人称360°)和Human360Occ(H3O,CARLA中人类视角360°,含RGB、深度、语义占位;标准划分城市内/跨城市)。OneOcc在QuadOcc上达到新最优,超越强视觉基线且媲美经典激光雷达基线;在H3O上,城市内提升+3.83 mIoU,跨城市提升+8.08。模块轻量,可部署于腿式/人形机器人实现全向感知。数据集与代码将公开于https://github.com/MasterHow/OneOcc。

原文摘要 · Abstract (English)

Robust 3D semantic occupancy is crucial for legged/humanoid robots, yet most semantic scene completion (SSC) systems target wheeled platforms with forward-facing sensors. We present OneOcc, a vision-only panoramic SSC framework designed for gait-introduced body jitter and 360° continuity. OneOcc combines: (i) Dual-Projection fusion (DP-ER) to exploit the annular panorama and its equirectangular unfolding, preserving 360° continuity and grid alignment; (ii) Bi-Grid Voxelization (BGV) to reason in Cartesian and cylindrical-polar spaces, reducing discretization bias and sharpening free/occupied boundaries; (iii) a lightweight decoder with Hierarchical AMoE-3D for dynamic multi-scale fusion and better long-range/occlusion reasoning; and (iv) plug-and-play Gait Displacement Compensation (GDC) learning feature-level motion correction without extra sensors. We also release two panoramic occupancy benchmarks: QuadOcc (real quadruped, first-person 360°) and Human360Occ (H3O) (CARLA human-ego 360° with RGB, Depth, semantic occupancy; standardized within-/cross-city splits). OneOcc sets a new state of the art on QuadOcc, outperforming strong vision baselines and remaining competitive with classical LiDAR baselines; on H3O it gains +3.83 mIoU (within-city) and +8.08 (cross-city). Modules are lightweight, enabling deployable full-surround perception for legged/humanoid robots. Datasets and code will be publicly available at https://github.com/MasterHow/OneOcc.

语义占位全景视觉腿式机器人轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。