arXiv:2602.05855cs.ROcs.LG2026-02

融合激光雷达与深度相机数据,用神经网络生成更精准的地形高度图。

A Hybrid Autoencoder for Robust Heightmap Generation from Fused Lidar and Depth Data for Humanoid Robot Locomotion

  • 用CNN+GRU混合结构提取空间与时间特征,提升地形感知鲁棒性。
  • 多模态融合使重建精度比纯深度或纯激光雷达提升7.2%和9.9%。
  • 引入3.2秒时序信息,有效减少地图漂移,适合复杂环境行走机器人。

在非结构化、以人为主的环境中部署类人机器人,可靠地形感知是关键前提。传统系统常依赖人工设计的单传感器流程,本文提出一种基于学习的框架,采用机器人中心的高度图表示。引入混合编码器-解码器结构(EDS),结合卷积神经网络(CNN)进行空间特征提取,以及门控循环单元(GRU)保持时间一致性。该架构融合了Intel RealSense深度相机、经高效球面投影处理的LIVOX MID-360激光雷达,以及机载惯性测量单元(IMU)的多模态数据。定量结果显示,多模态融合相较仅深度数据提升重建精度7.2%,相较仅激光雷达提升9.9%。此外,引入3.2秒的时序上下文可显著降低地图漂移。

原文摘要 · Abstract (English)

Reliable terrain perception is a critical prerequisite for the deployment of humanoid robots in unstructured, human-centric environments. While traditional systems often rely on manually engineered, single-sensor pipelines, this paper presents a learning-based framework that uses an intermediate, robot-centric heightmap representation. A hybrid Encoder-Decoder Structure (EDS) is introduced, utilizing a Convolutional Neural Network (CNN) for spatial feature extraction fused with a Gated Recurrent Unit (GRU) core for temporal consistency. The architecture integrates multimodal data from an Intel RealSense depth camera, a LIVOX MID-360 LiDAR processed via efficient spherical projection, and an onboard IMU. Quantitative results demonstrate that multimodal fusion improves reconstruction accuracy by 7.2% over depth-only and 9.9% over LiDAR-only configurations. Furthermore, the integration of a 3.2 s temporal context reduces mapping drift.

地形感知多模态融合类人机器人高度图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。