用10万小时真实机器人数据训练,让机器人听懂指令并快速适应新任务。
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

- 分两阶段训练:先用百万小时真实轨迹学通用动作,再对齐人类指令与机器人执行
- 在RoboCasa365上达成57.4%成功率,超越此前最优的46.6%
- 仅需少量数据即可高效微调,适合需要快速部署的机器人应用
我们提出Xiaomi-Robotics-1,一个基础视觉-语言-动作(VLA)模型,能够直接在未见过的环境中遵循多样化语言指令完成移动操作任务,并通过极少微调数据快速适应新任务。采用两阶段训练:预训练阶段利用超过10万小时的真实操作轨迹(通过UMI设备收集),结合可扩展的自动标注流水线,为轨迹片段生成描述场景状态变化的自然语言标签,赋予模型强泛化动作生成能力;后训练阶段对齐机器人本体与人类自然指令。大量实验表明,该模型具有显著的缩放效应:预训练数据量和模型规模越大,性能越优,且优势可直接传递至后训练,在未见环境中实现更强的实机表现。同时,该模型作为强基础策略,可在复杂灵巧任务中以高数据效率进行微调。在多个仿真基准测试中,Xiaomi-Robotics-1均优于现有方法,尤其在RoboCasa365上达到57.4%的成功率,超越此前最佳的46.6%;在RoboDojo上平均得分为20.07,显著高于先前的13.07。代码与模型检查点将公开。
原文摘要 · Abstract (English)
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.4% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。