arXiv:2410.23963cs.RO2024-10被引 14

用信息论理解人类示范,让机器人学会通用的手部操作技能

Exploiting Information Theory for Intuitive Robot Programming of Manual Activities

  • 基于香农信息论识别人机交互中的关键元素和信息流
  • 仅需一次示范即可生成机器人可执行的行为树,泛化能力强
  • 适合人机交互、具身智能领域研究者,开源数据集支持复现

观察学习为非编程专家用户提供了一种友好的技能迁移方式,模仿人类示范行为。现有方法多聚焦于轨迹复制,但在不同环境下的泛化能力受限。本文提出新框架,使机器人通过RGB视频理解手动任务的结构与目标,实现跨场景泛化。首次将香农信息论应用于手动任务分析,有效提取手物交互中的活跃场景元素,并量化其信息共享程度。结合场景图特性,以紧凑结构编码交互特征,将示范划分为任务块,用于生成机器人行为树。实验验证了该方法仅凭单一示范即可自动生成执行计划的有效性。同时发布开源数据集HANDSOME(多主体演示的手部技能数据集),推动该领域研究进展。

原文摘要 · Abstract (English)

Observational learning is a promising approach to enable people without expertise in programming to transfer skills to robots in a user-friendly manner, since it mirrors how humans learn new behaviors by observing others. Many existing methods focus on instructing robots to mimic human trajectories, but motion-level strategies often pose challenges in skills generalization across diverse environments. This paper proposes a novel framework that allows robots to achieve a higher-level understanding of human-demonstrated manual tasks recorded in RGB videos. By recognizing the task structure and goals, robots generalize what observed to unseen scenarios. We found our task representation on Shannon's Information Theory (IT), which is applied for the first time to manual tasks. IT helps extract the active scene elements and quantify the information shared between hands and objects. We exploit scene graph properties to encode the extracted interaction features in a compact structure and segment the demonstration into blocks, streamlining the generation of Behavior Trees for robot replicas. Experiments validated the effectiveness of IT to automatically generate robot execution plans from a single human demonstration. Additionally, we provide HANDSOME, an open-source dataset of HAND Skills demOnstrated by Multi-subjEcts, to promote further research and evaluation in this field.

机器人学习信息论行为树动作泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。