arXiv:2504.13351cs.ROcs.AI2025-04ICRA被引 9

用多模态数据教机器人从人类视频中学会操作任务。

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

  • 通过视频+肌电/音频信号融合,让模型逐步推理任务计划
  • 任务计划与控制参数提取准确率提升3倍,支持新任务泛化
  • 适合机器人学习、人机协作、具身智能研究者参考

从人类示范视频中学习执行操作任务是教导机器人的一种有前景的方法。然而,许多操作任务在执行过程中需要改变控制参数(如力度),仅靠视觉数据无法捕捉此类信息。本文利用臂带传感器测量肌肉活动、麦克风记录声音等传感设备,捕获人类操作过程中的细节,使机器人能够提取任务规划和控制参数以完成相同任务。为此,我们提出链式模态(Chain-of-Modality, CoM)提示策略,使视觉语言模型能够对多模态人类示范数据——视频结合肌电或音频信号——进行推理。通过逐模态整合信息,CoM不断优化任务计划并生成详细控制参数,使机器人仅需单个多模态人类视频即可执行操作任务。实验表明,与基线相比,CoM在任务计划与控制参数提取上准确率提升三倍,在真实机器人实验中对新任务设置和新物体均表现出强泛化能力。视频与代码已公开于 https://chain-of-modality.github.io

原文摘要 · Abstract (English)

Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during task execution, such as force, which visual data alone cannot capture. In this work, we leverage sensing devices such as armbands that measure human muscle activities and microphones that record sound, to capture the details in the human manipulation process, and enable robots to extract task plans and control parameters to perform the same task. To achieve this, we introduce Chain-of-Modality (CoM), a prompting strategy that enables Vision Language Models to reason about multimodal human demonstration data -- videos coupled with muscle or audio signals. By progressively integrating information from each modality, CoM refines a task plan and generates detailed control parameters, enabling robots to perform manipulation tasks based on a single multimodal human video prompt. Our experiments show that CoM delivers a threefold improvement in accuracy for extracting task plans and control parameters compared to baselines, with strong generalization to new task setups and objects in real-world robot experiments. Videos and code are available at https://chain-of-modality.github.io

机器人学习多模态视觉语言模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。