让机器人通过语音、动作等多方式持续学习,实时适应人类协作。
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
- 支持语音、手势、示范等多模态教学,实现多层次持续学习。
- 用户交互23小时,教出17个高层行为,监督减少22.1%,错误降低67.1%。
- 适合非专家用户快速上手,尤其适合创意协作如动画制作。
我们提出Vocal Sandbox框架,实现情境化人机协作的无缝衔接。系统具备从语音对话、物体关键点、体感示范等多种教学方式中,于多个抽象层次持续学习与自适应的能力。为此,我们设计轻量且可解释的学习算法,使用户能实时理解并协同调整机器人的能力。例如,在演示“绕物追踪”这一低层技能后,系统会可视化机器人对新物体的运动轨迹;用户也可通过语音指令,借助预训练语言模型组合低层技能,生成如“将物体收纳”等高层行为。我们在两个场景中评估:协作包装袋组装和LEGO定格动画。在前者中,8名非专家用户参与系统消融实验与用户研究,累计23小时交互,共教授17个高层行为,平均引入16个新低层技能,主动监督减少22.1%,自主性能提升19.7%,失败率下降67.1%。定性反馈显示,用户对易用性(+20.6%)与整体表现(+13.9%)高度认可。在后者中,经验用户与机器人连续协作两小时,逐步教学复杂动作,成功拍摄一段52秒(232帧)的定格动画。
原文摘要 · Abstract (English)
We introduce Vocal Sandbox, a framework for enabling seamless human-robot collaboration in situated environments. Systems in our framework are characterized by their ability to adapt and continually learn at multiple levels of abstraction from diverse teaching modalities such as spoken dialogue, object keypoints, and kinesthetic demonstrations. To enable such adaptation, we design lightweight and interpretable learning algorithms that allow users to build an understanding and co-adapt to a robot's capabilities in real-time, as they teach new behaviors. For example, after demonstrating a new low-level skill for "tracking around" an object, users are provided with trajectory visualizations of the robot's intended motion when asked to track a new object. Similarly, users teach high-level planning behaviors through spoken dialogue, using pretrained language models to synthesize behaviors such as "packing an object away" as compositions of low-level skills $-$ concepts that can be reused and built upon. We evaluate Vocal Sandbox in two settings: collaborative gift bag assembly and LEGO stop-motion animation. In the first setting, we run systematic ablations and user studies with 8 non-expert participants, highlighting the impact of multi-level teaching. Across 23 hours of total robot interaction time, users teach 17 new high-level behaviors with an average of 16 novel low-level skills, requiring 22.1% less active supervision compared to baselines and yielding more complex autonomous performance (+19.7%) with fewer failures (-67.1%). Qualitatively, users strongly prefer Vocal Sandbox systems due to their ease of use (+20.6%) and overall performance (+13.9%). Finally, we pair an experienced system-user with a robot to film a stop-motion animation; over two hours of continuous collaboration, the user teaches progressively more complex motion skills to shoot a 52 second (232 frame) movie.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。