用大模型驱动的增强现实界面,让宇航员远距离操控机器人更直观安全。
2024 NASA SUITS Report: LLM-Driven Immersive Augmented Reality User Interface for Robotics and Space Exploration
- 通过大模型语音控制和头戴AR设备实现无手操作
- 在无地面传感器条件下实现机器人实时6自由度定位
- 适合航天任务中的远程操控与工业巡检场景
随着计算技术进步,增强现实(AR)在叠加虚拟界面到物理对象方面取得进展,但复杂动态环境下的3D物体姿态估计仍是挑战。本项目针对移动AR中人机交互问题,提出URSA系统,用于应对未来阿耳忒弥斯任务等太空飞行需求。该系统整合头戴式AR设备(如HoloLens)、基于大语言模型的语音控制及机器人跟踪算法,实现空间感知与非侵入式交互。为提升精度,采用数字孪生定位技术,结合DTTD-Mobile数据集与ZED2相机,在噪声和遮挡环境下实现真实世界追踪。系统支持实时机器人控制与监控,无需依赖地面真值传感器,适用于高危或远程作业。关键贡献包括:(1)基于大模型语音输入的非侵入式AR界面;(2)面向非刚性机器人躯体的ZED2数据集;(3)用于任务可视化的本地任务控制台(LMCC);(4)优化深度融合的Transformer-based 6DoF姿态估计算法(DTTDNet);(5)端到端集成的宇航员任务支持系统。该研究推动数字孪生在机器人领域的应用,为航空航天与工业领域提供可扩展解决方案。
原文摘要 · Abstract (English)
As modern computing advances, new interaction paradigms have emerged, particularly in Augmented Reality (AR), which overlays virtual interfaces onto physical objects. This evolution poses challenges in machine perception, especially for tasks like 3D object pose estimation in complex, dynamic environments. Our project addresses critical issues in human-robot interaction within mobile AR, focusing on non-intrusive, spatially aware interfaces. We present URSA, an LLM-driven immersive AR system developed for NASA's 2023-2024 SUITS challenge, targeting future spaceflight needs such as the Artemis missions. URSA integrates three core technologies: a head-mounted AR device (e.g., HoloLens) for intuitive visual feedback, voice control powered by large language models for hands-free interaction, and robot tracking algorithms that enable accurate 3D localization in dynamic settings. To enhance precision, we leverage digital twin localization technologies, using datasets like DTTD-Mobile and specialized hardware such as the ZED2 camera for real-world tracking under noise and occlusion. Our system enables real-time robot control and monitoring via an AR interface, even in the absence of ground-truth sensors--vital for hazardous or remote operations. Key contributions include: (1) a non-intrusive AR interface with LLM-based voice input; (2) a ZED2-based dataset tailored for non-rigid robotic bodies; (3) a Local Mission Control Console (LMCC) for mission visualization; (4) a transformer-based 6DoF pose estimator (DTTDNet) optimized for depth fusion and real-time tracking; and (5) end-to-end integration for astronaut mission support. This work advances digital twin applications in robotics, offering scalable solutions for both aerospace and industrial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。