arXiv:2412.10410cs.AIcs.LG2024-12被引 13

用弱监督让机器人听懂多模态指令,不依赖大量标注数据

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents

  • 结合无标签演示与少量标注数据,通过隐变量模型学习行为
  • 在4个不同环境验证,能准确理解多模态指令并执行任务
  • 适合研究低成本可泛化智能体的开发者和机器人研究人员

开发能遵循多模态指令的智能体仍是机器人与人工智能领域的基础挑战。尽管在无标注数据集(无语言指令)上进行大规模预训练已使智能体学会多样化行为,但这些智能体常难以准确执行指令。虽然增加带指令标签的数据可缓解此问题,但在大规模下获取高质量标注极不现实。为此,我们将问题建模为半监督学习任务,提出GROOT-2——一种基于弱监督与隐变量模型的新方法训练的多模态可指令跟随智能体。该方法包含两个核心组件:约束自模仿,利用大量无标签示范让策略学习多样行为;人类意图对齐,使用少量带标注示范确保隐空间反映人类意图。GROOT-2在四个不同环境(涵盖视频游戏与机器人操作)中验证,展现出强大的多模态指令跟随能力。

原文摘要 · Abstract (English)

Developing agents that can follow multimodal instructions remains a fundamental challenge in robotics and AI. Although large-scale pre-training on unlabeled datasets (no language instruction) has enabled agents to learn diverse behaviors, these agents often struggle with following instructions. While augmenting the dataset with instruction labels can mitigate this issue, acquiring such high-quality annotations at scale is impractical. To address this issue, we frame the problem as a semi-supervised learning task and introduce GROOT-2, a multimodal instructable agent trained using a novel approach that combines weak supervision with latent variable models. Our method consists of two key components: constrained self-imitating, which utilizes large amounts of unlabeled demonstrations to enable the policy to learn diverse behaviors, and human intention alignment, which uses a smaller set of labeled demonstrations to ensure the latent space reflects human intentions. GROOT-2's effectiveness is validated across four diverse environments, ranging from video games to robotic manipulation, demonstrating its robust multimodal instruction-following capabilities.

多模态弱监督机器人指令跟随

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。