arXiv:2606.22971cs.ROcs.CV2026-06被引 2

构建面向人形机器人的全景立体占位数据集,支持真实-仿真-真实闭环训练。

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI

论文配图:Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI
图 1 · 摘自论文原文
  • 基于真实传感器参数生成物理精准的仿真数据,实现真实与仿真间闭环迭代。
  • 涵盖15个模拟室内场景和5个真实环境,共15.5万组带标注样本。
  • 模型在未见仿真与真实场景中均表现优异,适合人形机器人感知研究。

体素级占位预测对复杂环境中安全的机器人导航与交互至关重要。现有占位数据集多针对自动驾驶设计,存在以车辆为中心的视角偏差、远距离几何结构及静态道路先验等问题,难以适用于具身人形机器人感知。本文提出Humanoid-OmniOcc,一个大规模全景立体视觉占位数据集,专为人类形态机器人设计。数据集包含15个多样化模拟室内场景和5个真实环境,共产生超过15.5万组样本,覆盖广泛场景与风格多样性。关键创新在于采用Real2Sim2Real闭环范式:真实传感器规格驱动物理精准仿真,仿真生成大规模标注训练数据,模型在仿真中训练后直接在真实数据上评估,实现仿真到现实的持续优化。我们还提出人类形态环绕立体引导占位模型(Humanoid-OmniOcc),利用鲁棒深度先验实现高精度2D到3D的升维。大量实验表明,该数据集与模型显著优于单目基线,并在未见的模拟测试场景与真实环境上均表现出良好泛化能力,验证了Real2Sim2Real设计的有效性。代码与数据将在论文录用后公开于https://d-robotics-ai-lab.github.io/humanoid-omniocc。

原文摘要 · Abstract (English)

Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predominantly designed for autonomous driving with vehicle-centric biases -- forward-facing cameras, far-field geometry, and static road priors -- limiting their applicability to embodied humanoid perception. We present Humanoid-OmniOcc, a large-scale panoramic stereo-based occupancy dataset tailored for humanoid robots. The dataset encompasses 15 diverse simulated indoor scenes and 5 real-world environments, yielding over 155K samples with broad scene and style diversity. Importantly, the dataset is designed around a Real2Sim2Real closed-loop paradigm: real sensor specifications drive physically accurate simulation, simulation produces large-scale annotated training data, and models trained in simulation are directly evaluated on real-world captures -- enabling iterative refinement of the sim-to-real pipeline. We further propose \textbf{H}umanoid \textbf{S}urround \textbf{S}tereo-guided \textbf{Occ}upancy model (Humanoid-OmniOcc) that exploits robust depth priors for accurate 2D-to-3D lifting. Extensive experiments show that Humanoid-OmniOcc consistently outperforms monocular baselines and generalizes well to both unseen simulated test scenes and real-world environments, validating the effectiveness of the Real2Sim2Real design. Code and data will be available upon acceptance at https://d-robotics-ai-lab.github.io/humanoid-omniocc.

占位预测人形机器人仿真闭环立体视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。