arXiv:2606.17030cs.CV2026-06被引 7

用语言统一控制机器人世界,生成物理合理的未来视频。

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

论文配图:Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
图 1 · 摘自论文原文
  • 双流扩散模型融合视觉语言与视频潜空间,实现语言驱动的视频预测。
  • 构建860万条视频-文本数据集,覆盖20多种具身场景和500+动作类别。
  • 通用+专家渐进训练策略,支持零样本跨任务迁移与多视角一致性。

我们提出Qwen-RobotWorld,一种面向具身智能的语言条件视频世界模型。以自然语言为统一动作接口,该模型能从当前观测出发,预测机器人操作、自动驾驶、室内导航及人机任务转移等场景下的物理合理未来视觉轨迹。统一建模框架带来三大应用潜力:用于策略训练增强的合成数据生成、可扩展的虚拟环境评估,以及下游机器人控制的语言引导规划信号。实现方式包含三部分:a) 双流MMDiT结合冻结的Qwen2.5-VL语义与视频VAE潜变量,通过逐层联合注意力融合;b) 具身世界知识库(EWK),一个包含860万条视频-文本对(超2亿帧)的数据集,涵盖20余种具身形态与500多个动作类别;c) 通用+专家渐进式课程训练策略,先学习通用视觉先验,再在共享语言接口下注入具身专业化能力。大量实验表明其表现优异:在EWMBench和DreamGen Bench上排名第一,在WorldModelBench与PBench上超越所有开源模型。对RoboTwin-IF基准的零样本分析进一步验证其强泛化能力与多视角一致性。

原文摘要 · Abstract (English)

We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual trajectories from current observations across robotic manipulation, autonomous driving, indoor navigation, and human-to-robot transfer. This unified formulation provides three promising application directions: synthetic data generation for policy training augmentation, scalable virtual environments for policy evaluation, and language-guided planning signals for downstream robot control. This is achieved through a three-part design: a) Double-Stream MMDiT with MLLM Action Encoding, where a 60-layer double-stream diffusion transformer couples frozen Qwen2.5-VL semantics with video-VAE latents through layer-wise joint attention; b) Embodied World Knowledge (EWK), an 8.6M video-text corpus (200M+ frames) with action-language mapping over 20+ embodiments and 500+ action categories; and c) General+Expert Progressive Curriculum, a two-stage training strategy that first learns general visual priors and then injects embodied specialization under a shared language interface. Extensive results show strong competitiveness: ranks 1st overall on EWMBench and DreamGen Bench, outperforms all open-source models on WorldModelBench and PBench. Additional zero-shot analyses on RoboTwin-IF benchmark further support robust generalization and multi-view consistency.

具身智能视频生成语言控制机器人模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。