用单目摄像头实现城市人行道3D占位感知,兼顾精度与低成本。
Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning

- 融合激光雷达-图像对数据的几何信息与海量单目图像,实现混合2D-3D学习
- 在未配对单目图像上自动生成伪占位标签,避免昂贵3D标注
- 适用于配送机器人、轮椅等户外移动设备,对复杂人行道结构识别能力强
真实世界的人行道拥挤、杂乱且结构不规整,是移动机器人(如配送机器人、电动轮椅)安全导航的关键挑战。现有占位学习方法多针对道路自动驾驶设计,依赖大规模带配对激光雷达-图像数据集及密集3D监督信号,采集成本高,难以捕捉人行道特性。本文提出WalkOCC,一种面向人行道场景的单目3D占位感知框架。该方法通过配对序列生成伪占位监督信号,并在额外的仅图像数据上联合学习图像级表征,实现几何引导与可扩展学习的结合。无需3D占位标注即可获得稳定优化与良好泛化能力。大量实验表明,相比自监督图像基基线,其在预测精度、细微城市结构(如路缘、排水沟)分割以及环境与平台间迁移鲁棒性方面均有显著提升。为支持评估,我们构建了大尺度人行道感知数据集Sidewalk3D,包含多地点、多时段的激光雷达-相机配对序列及3D语义占位标注。代码与数据将公开。
原文摘要 · Abstract (English)
Sidewalks in the real world are crowded, cluttered, and less structured than roads, making 3D occupancy prediction a key ingredient for the safe navigation of mobile robots such as delivery bots and electric wheelchairs. Existing occupancy learning pipelines are largely designed for on-road autonomous driving and often train on large-scale paired LiDAR-RGB datasets with dense 3D supervision and multiple camera inputs, which are costly to collect and do not adequately capture sidewalk-specific characteristics. We propose WalkOCC, a hybrid Ray-marching monocular 3D occupancy perception framework for robots operating on sidewalks. WalkOCC explicitly couples geometric grounding from LiDAR-RGB paired data with scalable learning from large-scale unpaired monocular images. It bootstraps pseudo occupancy supervision from paired sequences and jointly learns image-level representations on additional 2D-only data. It yields stable optimization and improved generalization without requiring costly 3D occupancy annotations. Extensive experiments demonstrate consistent gains in prediction accuracy, fine-grained segmentation of subtle urban structures such as curbs and gutters, and robustness to environmental and cross-embodiment shifts compared with self-supervised image-based baselines. To facilitate evaluation and benchmarking, we also introduce Sidewalk3D, a large-scale sidewalk perception dataset with LiDAR-camera paired sequences collected across multiple locations and time periods, along with 3D semantic occupancy annotations for evaluation. Code and data will be made available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。