用语言指令控制实验室机器人,通过多层表示转换实现安全操作。
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation

- 将自然语言指令转化为可执行技能调用
- 通过传感器数据构建环境地图与物体位姿
- 提供调试接口,识别校准缺失等部署问题
开源机器人与基础模型降低了具身智能的门槛,但语言引导的实验室自动化仍需可靠地将指令与观测对齐为安全动作。本文报告了一个基于OpenArm的移动操作原型,集成双OpenArm机械臂、移动基座、垂直导轨、RGB-D感知、激光雷达建图、ROS2/MoveIt执行框架及配置化技能接口。系统围绕表示传递展开:自然语言请求被约束为注册技能调用,传感器观测被锚定为地图与物体位姿,物体先验提供角色与技能约束,运行时绑定将验证后的技能编译为可执行运动目标。通过预运行轨迹和启动检查评估该集成路径,揭示了校准缺失、物体资产不完整、真实场景视觉锚定未完成等显式部署障碍。这些中间表示作为实用的调试界面,帮助整合语言、感知、规划与机器人安全。
原文摘要 · Abstract (English)
Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable alignment from instructions and observations to safe actions. This field report presents an OpenArm-based mobile manipulation prototype for laboratory-style tasks, built by integrating dual OpenArm manipulators with a mobile base, vertical slide, RGB-D sensing, lidar-based mapping, ROS2/MoveIt execution, and profile-defined skill interfaces. The system is organized around representation handoffs: natural language requests are constrained into registered skill calls, sensor observations are grounded into maps and object poses, object priors provide role and skill constraints, and runtime bindings compile validated skills into executable motion goals. We use dry-run traces and startup checks to evaluate this integration path, showing how the prototype exposes missing calibration, incomplete object assets, and unfinished real-scene visual grounding as explicit deployment blockers. These intermediate representations serve as practical debugging interfaces for integrating language, perception, planning, and robot safety in embodied systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。