统一建模视觉、语言与动作,实现长时程具身智能体的端到端控制。
iFLYTEK-Embodied-Omni Technical Report

- 构建统一框架,通过共享注意力实现多模态信息协同。
- 在具身任务中实现高精度动作生成与未来状态预测。
- 适合研究具身智能、多模态大模型与机器人控制的学者。
通用具身智能体需理解多模态指令,预判环境演变,并在长时程内生成精确控制动作。现有方法多聚焦于视觉-语言推理、基于视频的世界建模或动作生成,而级联式管道(先生成未来观测再推断动作)易引入接口瓶颈并累积预测误差。本文提出 iFLYTEK-Embodied-Omni,一个统一的多模态基础模型,将视觉(视频与图像)、语言与动作整合于单一 Omni 框架中。其模态专用组件(视觉-语言模型、视频生成模型、动作生成模型)通过共享多模态自注意力进行通信,形成类脑-小脑协作机制:视觉-语言模型与视频生成模型构成高层‘大脑’,负责指令理解、任务规划、进度追踪与未来视觉状态预测;动作生成模型则作为低层‘小脑’,将规划子目标与共享多模态上下文直接转化为可执行的动作片段。为训练这些能力,我们融合人类示范与机器人交互的带动作标注和无动作标注的具身视频,结合具身推理、具身感知及通用图文数据构建综合性数据集。采用四阶段策略,逐步训练视觉-语言模型(VLM)、视频生成模型(VGM)与动作生成模型(AGM),最终联合微调完整模型。
原文摘要 · Abstract (English)
General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. Existing approaches typically specialize in visual-language reasoning, video-based world modeling, or action generation, while cascaded pipelines that first synthesize future observations and then infer actions can introduce interface bottlenecks and compound prediction errors. We present iFLYTEK-Embodied-Omni, a unified multimodal foundation model that jointly models vision(videos and images), language, and action within a single Omni framework. Its modality-specific visual-language, video-generation, and action-generation components communicate through shared multimodal self-attention. This design establishes brain-cerebellum collaboration: the vision-language modeland video generation model form a high-level brain for instruction understanding, task planning, progress tracking, and future visual-state prediction, whereas the action generation modelserves as a low-level cerebellum that directly converts planned subgoals and shared multimodal context into executable action chunks. To develop these capabilities, we combine action-annotated and action-free embodied videos from human demonstrations and robot interactions with embodied reasoning, embodied perception, and general-purpose image-text data to construct a comprehensive dataset. We further adopt a four-stage strategy that progressively trains the VLM, VGM, and AGM before jointly fine-tuning the complete model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。