arXiv:2504.15369cs.LGcs.AI2025-04ICLR被引 17

用少量机器人数据适配大模型,让机器人通过文字指令学会新任务。

Solving New Tasks by Adapting Internet Video Knowledge

  • 结合互联网视频大模型与少量机器人实拍数据进行适配。
  • 仅需少量示例数据即可实现新任务的文本驱动泛化。
  • 新方法对数据质量不敏感,即使低质量演示也能成功完成任务。

视频生成模型在机器人领域展现巨大潜力,可作为视觉规划器或策略监督者。当在互联网规模数据上预训练时,这些模型能深刻理解自然语言对齐,从而通过文本条件实现对新下游行为的泛化。然而,它们可能对特定环境细节不敏感。相反,在机器人行为的域内数据上训练视频模型可自然编码环境特异性,但可用示范数据量有限,难以支持通过自然语言指定实现对未见任务的泛化。本文研究了不同融合域内信息与大规模预训练视频模型的适配技术,探索其在实现文本条件泛化方面的效果,并考虑各自独立的数据与资源需求。实验表明,使用小规模示例数据适配强大视频模型,可有效促进新行为的泛化。特别地,我们提出一种新适配策略——逆概率适配(Inverse Probabilistic Adaptation),该方法在多种机器人任务和设置中均表现出色,且对适配数据质量具有鲁棒性,即使仅有次优的域内演示,也能成功解决新任务。

原文摘要 · Abstract (English)

Video generative models demonstrate great promise in robotics by serving as visual planners or as policy supervisors. When pretrained on internet-scale data, such video models intimately understand alignment with natural language, and can thus facilitate generalization to novel downstream behavior through text-conditioning. However, they may not be sensitive to the specificities of the particular environment the agent inhabits. On the other hand, training video models on in-domain examples of robotic behavior naturally encodes environment-specific intricacies, but the scale of available demonstrations may not be sufficient to support generalization to unseen tasks via natural language specification. In this work, we investigate different adaptation techniques that integrate in-domain information with large-scale pretrained video models, and explore the extent to which they enable novel text-conditioned generalization for robotic tasks, while also considering their independent data and resource considerations. We successfully demonstrate across robotic environments that adapting powerful video models with small scales of example data can successfully facilitate generalization to novel behaviors. In particular, we present a novel adaptation strategy, termed Inverse Probabilistic Adaptation, that not only consistently achieves strong generalization performance across robotic tasks and settings, but also exhibits robustness to the quality of adaptation data, successfully solving novel tasks even when only suboptimal in-domain demonstrations are available.

视频生成机器人文本条件小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。