arXiv:2511.01914cs.CVcs.AI2025-11

iFlyBot-VLA用双层动作表示提升机器人视觉语言决策能力

iFlyBot-VLA Technical Report

  • 用跨体感操控数据预训练隐式动作模型,融合高阶意图与低阶动态
  • 在LIBERO Franka上实现92.3%成功率,实测表现优于主流方法
  • 适合做具身智能、机器人操作与多模态决策的研究者参考

我们提出iFlyBot-VLA,一种基于新型框架的大规模视觉-语言-动作(VLA)模型。主要贡献包括:(1) 在大规模人类与机器人操作视频上训练的隐式动作模型;(2) 双层次动作表示框架,联合监督视觉-语言模型(VLM)与动作专家;(3) 混合训练策略,结合机器人轨迹数据与通用问答及空间问答数据集,有效增强VLM骨干的3D感知与推理能力。具体地,VLM被训练预测两种互补动作形式:源自跨体感操作数据预训练的隐式动作,捕捉高层意图;通过连续控制信号频域变换获得的结构化离散动作标记,编码底层动态。该双重监督使语言、视觉与动作表征空间对齐,支持VLM直接生成动作。在LIBERO Franka基准上的实验结果表明本框架具有优势,真实世界评估进一步显示iFlyBot-VLA在多样且复杂的操作任务中达到有竞争力的成功率。此外,我们将开源部分自建数据集以支持社区后续研究。

原文摘要 · Abstract (English)

We introduce iFlyBot-VLA, a large-scale Vision-Language-Action (VLA) model trained under a novel framework. The main contributions are listed as follows: (1) a latent action model thoroughly trained on large-scale human and robotic manipulation videos; (2) a dual-level action representation framework that jointly supervises both the Vision-Language Model (VLM) and the action expert during training; (3) a mixed training strategy that combines robot trajectory data with general QA and spatial QA datasets, effectively enhancing the 3D perceptual and reasoning capabilities of the VLM backbone. Specifically, the VLM is trained to predict two complementary forms of actions: latent actions, derived from our latent action model pretrained on cross-embodiment manipulation data, which capture implicit high-level intentions; and structured discrete action tokens, obtained through frequency-domain transformations of continuous control signals, which encode explicit low-level dynamics. This dual supervision aligns the representation spaces of language, vision, and action, enabling the VLM to directly contribute to action generation. Experimental results on the LIBERO Franka benchmark demonstrate the superiority of our frame-work, while real-world evaluations further show that iFlyBot-VLA achieves competitive success rates across diverse and challenging manipulation tasks. Furthermore, we plan to open-source a portion of our self-constructed dataset to support future research in the community

机器人操作视觉语言动作多模态学习具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。