arXiv:2606.17846cs.ROcs.CV2026-06被引 31

用对齐框架让机器人操控模型规模突破,实现跨平台零样本泛化。

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

论文配图:Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
图 1 · 摘自论文原文
  • 构建视觉-语言-动作统一对齐框架,支持多源数据大规模协同训练。
  • 基于开源数据生成3.8万小时预训练语料,实现零样本指令执行与抗扰动能力。
  • 适用于真实机器人平台,适合研究通用机器人智能的开发者和工程师。

语言与多模态基础模型通过在统一框架下对齐异构数据并规模化训练,实现了强泛化能力。本文探讨这一扩展范式能否应用于机器人操控任务以实现真正泛化。由于操控数据天然异构、采集成本高且多样性有限,对齐与规模化同步实现极具挑战。我们提出Qwen-RobotManip,一个基于Qwen-VL的通用视觉-语言-动作基础模型。该模型引入跨表征、运动与行为维度的统一对齐框架,使大规模多源训练变得连贯而非冲突。此对齐能力使得模型可吸收远超以往训练模式的数据规模。通过人到机器的合成流水线,将15种平台的自我视角手部示范转化为机器人轨迹,并通过严格清洗流程调和异构数据集。仅使用开源数据与人类视频,未依赖专有数据采集,构建约38,100小时的预训练语料库。模型展现出涌现的泛化能力:零样本指令遵循、对扰动的鲁棒性、反应式错误恢复及跨机体迁移。标准基准无法有效评估预训练质量,因此采用包括RoboCasa365、LIBERO-Plus、EBench、RoboTwin-Clean2Rand、RoboTwin-IF和RoboTwin-XE在内的域外测试设置。Qwen-RobotManip在所有域外设置中显著优于先前最先进模型(如π0.5),在RoboChallenge中排名第一,相对提升达20%,并在AgileX ALOHA、Franka、UR、ARX等真实机器人平台上得到验证。

原文摘要 · Abstract (English)

Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $π$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.

机器人操控基础模型对齐训练泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。