arXiv:2603.05355cs.RO2026-03

用360°激光雷达实现无需移动的全向操作,让人形机器人在大空间中稳定抓取。

OmniDP: Beyond-FOV Large-Workspace Humanoid Manipulation with Omnidirectional 3D Perception

  • 基于激光雷达点云,通过时间感知注意力池化实现全景3D感知
  • 在模拟与真实场景中,大空间杂乱环境下的成功率超基线方法27%以上
  • 适合需要大范围精准操作的工业搬运、救援等场景

人形机器人在非结构化环境中进行灵巧操作仍受感知能力限制,有效工作区受限。当物理条件无法移动机器人时,全向感知比颜色或语义信息更为关键。尽管视觉-运动策略学习取得进展,传统RGB-D方案存在视野狭窄和自遮挡问题,需频繁移动基座,带来运动不确定性与安全风险。现有扩展感知方法如主动视觉系统和第三视角相机则引入机械复杂性、标定依赖和延迟,影响实时可靠性。本文提出OmniDP,一种端到端的激光雷达驱动3D视觉-运动策略,可在大工作区实现鲁棒操作。该方法通过时间感知注意力池化处理全景点云,高效编码稀疏3D数据并捕捉时间依赖性。360°感知使机器人无需频繁重定位即可交互广阔区域内的物体。为支持策略学习,我们构建了全身遥操作系统,用于高效采集全身协同数据。大量仿真与真实环境实验表明,OmniDP在大空间及杂乱场景中表现稳健,优于依赖中心视角深度相机的基线方法。

原文摘要 · Abstract (English)

The deployment of humanoid robots for dexterous manipulation in unstructured environments remains challenging due to perceptual limitations that constrain the effective workspace. In scenarios where physical constraints prevent the robot from repositioning itself, maintaining omnidirectional awareness becomes far more critical than color or semantic information.While recent advances in visuomotor policy learning have improved manipulation capabilities, conventional RGB-D solutions suffer from narrow fields of view (FOV) and self-occlusion, requiring frequent base movements that introduce motion uncertainty and safety risks. Existing approaches to expanding perception, including active vision systems and third-view cameras, introduce mechanical complexity, calibration dependencies, and latency that hinder reliable real-time performance. In this work, We propose OmniDP, an end-to-end LiDAR-driven 3D visuomotor policy that enables robust manipulation in large workspaces. Our method processes panoramic point clouds through a Time-Aware Attention Pooling mechanism, efficiently encoding sparse 3D data while capturing temporal dependencies. This 360° perception allows the robot to interact with objects across wide areas without frequent repositioning. To support policy learning, we develop a whole-body teleoperation system for efficient data collection on full-body coordination. Extensive experiments in simulation and real-world environments show that OmniDP achieves robust performance in large-workspace and cluttered scenarios, outperforming baselines that rely on egocentric depth cameras.

人形机器人3D感知大空间操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。