无需标记物,用普通摄像头实现多机器人实时位姿共享。
Zero-Splat TeleAssist: A Zero-Shot Pose Estimation Framework for Semantic Teleoperation
- 融合视觉语言分割与单目深度,构建3D世界模型
- 通过加权主成分分析提取6自由度位姿,支持实时交互
- 零样本适配,适合无深度传感器的远程操作场景
我们提出Zero-Splat TeleAssist,一种零样本传感器融合管道,可将普通CCTV视频流转化为多边远程操作中的共享6-自由度世界模型。通过整合视觉-语言分割、单目深度估计、加权主成分分析(weighted-PCA)位姿提取以及3D高斯溅射(3DGS),TeleAssist在无标记物、无深度传感器的以交互为中心的操作设置下,为每位操作员提供多个机器人的实时全局位置与姿态。该框架实现了跨设备的协同感知与定位,显著降低了远程操作系统的部署门槛。
原文摘要 · Abstract (English)
We introduce Zero-Splat TeleAssist, a zero-shot sensor-fusion pipeline that transforms commodity CCTV streams into a shared, 6-DoF world model for multilateral teleoperation. By integrating vision-language segmentation, monocular depth, weighted-PCA pose extraction, and 3D Gaussian Splatting (3DGS), TeleAssist provides every operator with real-time global positions and orientations of multiple robots without fiducials or depth sensors in an interaction-centric teleoperation setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。