arXiv:2503.09010cs.RO2025-03被引 19

用全景视觉与激光雷达融合,让机器人看清360度环境

HumanoidPano: Hybrid Spherical Panoramic-LiDAR Cross-Modal Perception for Humanoid Robots

  • 通过球面变换实现全景图与激光雷达的精准对齐
  • 在360BEV-Matterport上达到当前最佳性能
  • 适合需要全方位感知的仿人机器人应用

仿人机器人因结构限制存在严重自遮挡和视场受限问题。本文提出HumanoidPano,一种融合全景视觉与激光雷达的混合跨模态感知框架。通过球面视觉变压器实现几何感知对齐,将360°视觉上下文与激光雷达精确深度测量无缝融合。首先,球面几何约束(SGC)利用全景相机射线特性,指导畸变正则化的采样偏移以实现几何对齐;其次,空间可变形注意力(SDA)通过球面偏移聚合分层3D特征,实现高效360°到鸟瞰图(BEV)融合,生成几何完整的物体表征;第三,全景增强(AUG)结合跨视角变换与语义对齐,提升数据增强时BEV-全景特征的一致性。大量实验表明,该方法在360BEV-Matterport基准上表现领先。真实机器人平台部署验证了系统在复杂环境中通过全景-LiDAR协同感知生成准确的鸟瞰图分割图,直接支持下游导航任务。本工作为仿人机器人具身感知建立了新范式。

原文摘要 · Abstract (English)

The perceptual system design for humanoid robots poses unique challenges due to inherent structural constraints that cause severe self-occlusion and limited field-of-view (FOV). We present HumanoidPano, a novel hybrid cross-modal perception framework that synergistically integrates panoramic vision and LiDAR sensing to overcome these limitations. Unlike conventional robot perception systems that rely on monocular cameras or standard multi-sensor configurations, our method establishes geometrically-aware modality alignment through a spherical vision transformer, enabling seamless fusion of 360 visual context with LiDAR's precise depth measurements. First, Spherical Geometry-aware Constraints (SGC) leverage panoramic camera ray properties to guide distortion-regularized sampling offsets for geometric alignment. Second, Spatial Deformable Attention (SDA) aggregates hierarchical 3D features via spherical offsets, enabling efficient 360°-to-BEV fusion with geometrically complete object representations. Third, Panoramic Augmentation (AUG) combines cross-view transformations and semantic alignment to enhance BEV-panoramic feature consistency during data augmentation. Extensive evaluations demonstrate state-of-the-art performance on the 360BEV-Matterport benchmark. Real-world deployment on humanoid platforms validates the system's capability to generate accurate BEV segmentation maps through panoramic-LiDAR co-perception, directly enabling downstream navigation tasks in complex environments. Our work establishes a new paradigm for embodied perception in humanoid robotics.

仿人机器人多模态感知全景视觉激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。