六摄像头加LiDAR实现机器人360度实时视觉,减少眩晕感。
RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception

- 六摄像头融合LiDAR生成全景360°视图,支持实时渲染与重建。
- 通过球坐标系与分层体素先验,仅用少量高精度3D高斯表示场景。
- 适用于远程操控、数据采集等需沉浸式视觉的机器人任务。
环绕视图感知在机器人导航与人机协同操作中日益重要,但现有系统多限于窄视角或需手动切换多相机,导致头戴设备中运动抖动引发仿真眩晕。本文提出一种结合六摄像头与LiDAR的360°全景视觉系统,满足具身部署的几何与实时性要求。进一步设计了 extsc{RobotPan}前馈框架,从校准的稀疏视角输入预测度量尺度且紧凑的3D高斯,实现实时渲染、重建与流传输。该方法将多视角特征映射至统一球坐标系,利用分层球体素先验,在机器人附近分配高分辨率,远距离降低分辨率,有效减少计算冗余而不牺牲精度。为支持长序列,采用在线融合策略更新动态内容,同时选择性更新外观以防止静态区域无限膨胀。我们还发布了面向机器人360°新视角合成与度量3D重建的多传感器数据集,涵盖真实平台上的导航、操作与移动任务。实验表明, extsc{RobotPan}在重建质量上优于现有前馈方法,同时生成的高斯数量显著更少,可实现实用化的实时具身部署。
原文摘要 · Abstract (English)
Surround-view perception is increasingly important for robotic navigation and loco-manipulation, especially in human-in-the-loop settings such as teleoperation, data collection, and emergency takeover. However, current robotic visual interfaces are often limited to narrow forward-facing views, or, when multiple on-board cameras are available, require cumbersome manual switching that interrupts the operator's workflow. Both configurations suffer from motion-induced jitter that causes simulator sickness in head-mounted displays. We introduce a surround-view robotic vision system that combines six cameras with LiDAR to provide full 360$^\circ$ visual coverage, while meeting the geometric and real-time constraints of embodied deployment. We further present \textsc{RobotPan}, a feed-forward framework that predicts \emph{metric-scaled} and \emph{compact} 3D Gaussians from calibrated sparse-view inputs for real-time rendering, reconstruction, and streaming. \textsc{RobotPan} lifts multi-view features into a unified spherical coordinate representation and decodes Gaussians using hierarchical spherical voxel priors, allocating fine resolution near the robot and coarser resolution at larger radii to reduce computational redundancy without sacrificing fidelity. To support long sequences, our online fusion updates dynamic content while preventing unbounded growth in static regions by selectively updating appearance. Finally, we release a multi-sensor dataset tailored to 360$^\circ$ novel view synthesis and metric 3D reconstruction for robotics, covering navigation, manipulation, and locomotion on real platforms. Experiments show that \textsc{RobotPan} achieves competitive quality against prior feed-forward reconstruction and view-synthesis methods while producing substantially fewer Gaussians, enabling practical real-time embodied deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。