通过多模态演示学习,让机器人理解人类动作意图并实现对齐
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
- 用视觉与体素化深度数据联合建模人类与机器人动作
- 在10个场景下5名用户数据上,人机模型准确率均超71%
- 适合研究人机协作与模仿学习的开发者参考
在非结构化环境中,理解人类与机器人动作之间的对应关系对评估决策对齐至关重要。本文提出一种多模态演示学习框架,显式地从RGB视频中建模人类示范,并在体素化的RGB-D空间中建模机器人示范。以RH20T数据集中的“拾取与放置”任务为例,利用5名用户在10个不同场景下的数据进行训练。方法结合基于ResNet的视觉编码器用于人类意图建模,以及基于Perceiver Transformer的体素化机器人动作预测。经过2000轮训练后,人类模型达到71.67%的准确率,机器人模型达到71.8%的准确率,证明了该框架在复杂操作任务中对齐多模态人机行为的潜力。
原文摘要 · Abstract (English)
Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. We propose a multimodal demonstration learning framework that explicitly models human demonstrations from RGB video with robot demonstrations in voxelized RGB-D space. Focusing on the "pick and place" task from the RH20T dataset, we utilize data from 5 users across 10 diverse scenes. Our approach combines ResNet-based visual encoding for human intention modeling and a Perceiver Transformer for voxel-based robot action prediction. After 2000 training epochs, the human model reaches 71.67% accuracy, and the robot model achieves 71.8% accuracy, demonstrating the framework's potential for aligning complex, multimodal human and robot behaviors in manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。