让机器人听懂自然语言指令并生成类人运动轨迹
Speech-to-Trajectory: Learning Human-Like Verbal Guidance for Robot Motion
- 用行为克隆学习人类引导的运动数据,直接映射语音到动作
- 通过GPT生成语义增广命令,提升对不同说法的适应性
- 无需提示工程,适合真实场景下多平台机器人部署
将机器人融入实际应用需要其能理解并执行未经训练用户发出的自然语言指令。由于人类语言的固有差异,相同指令可能表述不同,但机器人应保持一致行为。尽管大语言模型(LLMs)提升了语言理解能力,却常因用户表达变化而失效,依赖预设命令且输出不可预测。本文提出指令语言模型(DLM),一种直接将口语指令映射为可执行运动轨迹的新型框架,避免使用预定义短语。DLM 在模拟的人类引导机器人运动示范上使用行为克隆(BC)进行训练,并引入基于GPT的语义增强技术,生成多样化的训练指令变体,均标注同一运动轨迹。DLM 进一步结合基于扩散策略的轨迹生成方法,实现自适应运动优化与随机采样。相比基于LLM的方法,DLM确保了稳定可预测的动作输出,无需大量提示工程,支持实时机器人引导。由于DLM从轨迹数据中学习,具备本体无关性,可在多种机器人平台上部署。实验表明,DLM在指令泛化能力、减少对结构化表达依赖方面表现更优,并实现了类人运动效果。
原文摘要 · Abstract (English)
Full integration of robots into real-life applications necessitates their ability to interpret and execute natural language directives from untrained users. Given the inherent variability in human language, equivalent directives may be phrased differently, yet require consistent robot behavior. While Large Language Models (LLMs) have advanced language understanding, they often falter in handling user phrasing variability, rely on predefined commands, and exhibit unpredictable outputs. This letter introduces the Directive Language Model (DLM), a novel speech-to-trajectory framework that directly maps verbal commands to executable motion trajectories, bypassing predefined phrases. DLM utilizes Behavior Cloning (BC) on simulated demonstrations of human-guided robot motion. To enhance generalization, GPT-based semantic augmentation generates diverse paraphrases of training commands, labeled with the same motion trajectory. DLM further incorporates a diffusion policy-based trajectory generation for adaptive motion refinement and stochastic sampling. In contrast to LLM-based methods, DLM ensures consistent, predictable motion without extensive prompt engineering, facilitating real-time robotic guidance. As DLM learns from trajectory data, it is embodiment-agnostic, enabling deployment across diverse robotic platforms. Experimental results demonstrate DLM's improved command generalization, reduced dependence on structured phrasing, and achievement of human-like motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。