arXiv:2607.07675cs.CV2026-07被引 7

用专家混合模型打造面向机器人智能的视频预训练模型

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

论文配图:Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
图 1 · 摘自论文原文
  • 采用专家混合架构提升建模能力与推理效率的平衡
  • 在多维度奖励机制下实现物理合理性与任务完成度的对齐
  • 专为机器人操作设计,适合具身智能研究者使用

尽管机器人控制领域前景广阔,视频生成模型仍因过度关注内容创作而存在领域不匹配问题。其设计往往优先考虑视觉保真度和创意性,而非计算效率与物理真实性。本文提出LingBot-Video,一种基于DiT的视频预训练范式,专为具身智能设计。从架构看,采用专家混合(MoE)而非密集结构,实现建模能力与推理效率的更好权衡,并成功从零开始规模化;从数据看,构建数据画像引擎,将标准互联网视频与大量机器人相关影像(含操作、导航、第一视角)融合,赋予模型对动作与世界动态的内在理解;从训练看,开发多维奖励系统,强化物理合理性与任务完成度,超越传统美学、提示遵循与运动一致性标准。全面评估验证了其作为视频基础模型的性能与效率。我们开源LingBot-Video,作为首个大规模、开源的MoE视频基础模型,开创性地弥合数字创造力与物理执行之间的鸿沟。

原文摘要 · Abstract (English)

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

视频生成专家混合具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。