用流模型预测动作,让双臂机器人更高效地学会复杂操作。
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

- 分两阶段用光流衔接文本与视频,精准理解语言指令
- 仅需少量真实数据,就能在仿真和真实机器人上成功执行任务
- 适合想低成本训练双臂机器人的研究者和工程师
由于动作空间大且需双臂协同,学习通用的双臂操作策略对具身智能体极具挑战。现有方法依赖视觉-语言-动作(VLA)模型获取双臂策略,但单臂数据或预训练VLA模型的知识迁移常因双臂数据稀缺及单双臂操作本质差异而失效。本文提出一种新型双臂基础策略:微调先进文到视频模型以预测机器人轨迹,并训练轻量级扩散策略生成动作。鉴于文到视频模型缺乏具身知识,我们设计两阶段范式,从预训练文到视频模型中分离出独立的文到光流与光流到视频模型。光学流作为图像间细微运动的紧凑表示,文到光流模型将其语义化,光流到视频模型则据此进行细粒度视频预测。该方法缓解了单阶段文到视频预测中的语言模糊性,显著降低对机器人低级动作数据的需求。实验中,我们采集高质量双臂机器人操作数据,在仿真与真实场景下均验证了方法的有效性。
原文摘要 · Abstract (English)
Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existing approaches rely on Vision-Language-Action (VLA) models to acquire bimanual policies. However, transferring knowledge from single-arm datasets or pre-trained VLA models often fails to generalize effectively, primarily due to the scarcity of bimanual data and the fundamental differences between single-arm and bimanual manipulation. In this paper, we propose a novel bimanual foundation policy by fine-tuning the leading text-to-video models to predict robot trajectories and training a lightweight diffusion policy for action generation. Given the lack of embodied knowledge in text-to-video models, we introduce a two-stage paradigm that fine-tunes independent text-to-flow and flow-to-video models derived from a pre-trained text-to-video model. Specifically, optical flow serves as an intermediate variable, providing a concise representation of subtle movements between images. The text-to-flow model predicts optical flow to concretize the intent of language instructions, and the flow-to-video model leverages this flow for fine-grained video prediction. Our method mitigates the ambiguity of language in single-stage text-to-video prediction and significantly reduces the robot-data requirement by avoiding direct use of low-level actions. In experiments, we collect high-quality manipulation data for real dual-arm robot, and the results of simulation and real-world experiments demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。