让机器人看懂环境并自主执行任务,跨平台通用视觉语言模型。
iFlyBot-VLM Technical Report
- 将视觉信息抽象为通用操作语言,实现感知与控制闭环
- 在10个主流数据集上表现最优,保持强泛化能力
- 适合研究通用机器人智能的学者与开发者
我们提出iFlyBot-VLM,一个用于提升具身智能领域的通用视觉语言模型。其核心目标是弥合高维环境感知与低层机器人运动控制之间的跨模态语义鸿沟。该模型将复杂的视觉与空间信息抽象为与身体无关、可迁移的运行语言,从而实现不同机器人平台间的无缝感知-动作闭环协同。iFlyBot-VLM架构系统设计了四项关键功能:1)空间理解与度量推理;2)交互目标定位;3)动作抽象与控制参数生成;4)任务规划与技能编排。我们将其视为具身人工智能的可扩展基础模型,推动从专用任务系统向通用认知智能体演进。在10个主流具身智能相关VLM基准数据集(如Blink、Where2Place)上进行评估,性能最优且保持良好泛化性。我们将公开训练数据与模型权重,以促进具身智能领域的发展。
原文摘要 · Abstract (English)
We introduce iFlyBot-VLM, a general-purpose Vision-Language Model (VLM) used to improve the domain of Embodied Intelligence. The central objective of iFlyBot-VLM is to bridge the cross-modal semantic gap between high-dimensional environmental perception and low-level robotic motion control. To this end, the model abstracts complex visual and spatial information into a body-agnostic and transferable Operational Language, thereby enabling seamless perception-action closed-loop coordination across diverse robotic platforms. The architecture of iFlyBot-VLM is systematically designed to realize four key functional capabilities essential for embodied intelligence: 1) Spatial Understanding and Metric Reasoning; 2) Interactive Target Grounding; 3) Action Abstraction and Control Parameter Generation; 4) Task Planning and Skill Sequencing. We envision iFlyBot-VLM as a scalable and generalizable foundation model for embodied AI, facilitating the progression from specialized task-oriented systems toward generalist, cognitively capable agents. We conducted evaluations on 10 current mainstream embodied intelligence-related VLM benchmark datasets, such as Blink and Where2Place, and achieved optimal performance while preserving the model's general capabilities. We will publicly release both the training data and model weights to foster further research and development in the field of Embodied Intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。