arXiv:2504.11493cs.ROcs.AI2025-04

通过多模态演示学习,让机器人理解人类动作意图并实现对齐

Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning

  • 用视觉与体素化深度数据联合建模人类与机器人动作
  • 在10个场景下5名用户数据上,人机模型准确率均超71%
  • 适合研究人机协作与模仿学习的开发者参考

在非结构化环境中,理解人类与机器人动作之间的对应关系对评估决策对齐至关重要。本文提出一种多模态演示学习框架,显式地从RGB视频中建模人类示范,并在体素化的RGB-D空间中建模机器人示范。以RH20T数据集中的“拾取与放置”任务为例,利用5名用户在10个不同场景下的数据进行训练。方法结合基于ResNet的视觉编码器用于人类意图建模,以及基于Perceiver Transformer的体素化机器人动作预测。经过2000轮训练后,人类模型达到71.67%的准确率,机器人模型达到71.8%的准确率,证明了该框架在复杂操作任务中对齐多模态人机行为的潜力。

原文摘要 · Abstract (English)

Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. We propose a multimodal demonstration learning framework that explicitly models human demonstrations from RGB video with robot demonstrations in voxelized RGB-D space. Focusing on the "pick and place" task from the RH20T dataset, we utilize data from 5 users across 10 diverse scenes. Our approach combines ResNet-based visual encoding for human intention modeling and a Perceiver Transformer for voxel-based robot action prediction. After 2000 training epochs, the human model reaches 71.67% accuracy, and the robot model achieves 71.8% accuracy, demonstrating the framework's potential for aligning complex, multimodal human and robot behaviors in manipulation tasks.

人机协作模仿学习多模态动作对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。