用眼神+语音控制机器人,实现直观的远程力控操作。
Visio-Verbal Teleimpedance Interface: Enabling Semi-Autonomous Control of Physical Interaction via Eye Tracking and Speech
- 结合眼动追踪与语音指令,通过视觉语言模型理解操作意图。
- 在滑槽任务中成功实现对机械臂3D刚度椭球的精准调控。
- 适合人机协作、远程操控等需要自然交互的工业场景。
本文提出一种基于视觉-语音的遥操作阻抗接口,通过融合操作员的眼动与语音交互,远程控制机器人的3D刚度椭球。眼动由Tobii Pro Glasses 2设备追踪,识别操作员关注的场景区域;语音指令则通过GPT-4o驱动的视觉语言模型(VLM)处理,以理解意图或提供修正。系统据此生成对应物理交互动作的刚度矩阵。实验在包含Force Dimension Sigma.7力反馈设备与Kuka LBR iiwa机械臂的平台上进行。首先优化了接口的提示词配置;随后在滑槽插入任务中验证了该接口的多种功能,证明其能实现高效、直观的远程力控交互。
原文摘要 · Abstract (English)
The paper presents a visio-verbal teleimpedance interface for commanding 3D stiffness ellipsoids to the remote robot with a combination of the operator's gaze and verbal interaction. The gaze is detected by an eye-tracker, allowing the system to understand the context in terms of what the operator is currently looking at in the scene. Along with verbal interaction, a Visual Language Model (VLM) processes this information, enabling the operator to communicate their intended action or provide corrections. Based on these inputs, the interface can then generate appropriate stiffness matrices for different physical interaction actions. To validate the proposed visio-verbal teleimpedance interface, we conducted a series of experiments on a setup including a Force Dimension Sigma.7 haptic device to control the motion of the remote Kuka LBR iiwa robotic arm. The human operator's gaze is tracked by Tobii Pro Glasses 2, while human verbal commands are processed by a VLM using GPT-4o. The first experiment explored the optimal prompt configuration for the interface. The second and third experiments demonstrated different functionalities of the interface on a slide-in-the-groove task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。