arXiv:2508.20840cs.ROcs.AI2025-08被引 5

用基础动作建模机器人世界,提升学习效率与泛化能力

Learning Primitive Embodied World Models: Towards Scalable Robotic Learning

  • 限定短时视频生成,聚焦基础动作单元
  • 实现语言与动作的细粒度对齐,降低训练复杂度
  • 适合需要可解释、可组合控制的机器人任务

尽管基于视频生成的具身世界模型受到关注,但其对大规模具身交互数据的依赖仍是主要瓶颈。具身数据稀缺、采集困难且维度高,导致语言与动作对齐粒度不足,加剧了长时序视频生成难度,阻碍生成模型在具身领域实现类似GPT的突破。我们观察到:具身数据多样性远超有限的基础动作空间。基于此,提出新范式——基础动作具身世界模型(PEWM)。通过限制视频生成在固定短时程内,该方法实现:1)语言概念与机器人动作视觉表征的细粒度对齐;2)降低学习复杂度;3)提升具身数据收集的数据效率;4)减少推理延迟。结合模块化视觉-语言模型(VLM)规划器与起点-终点热力图引导机制(SGG),PEWM支持灵活闭环控制,并实现基础策略在复杂任务上的组合泛化。框架利用视频模型中的时空视觉先验与VLM的语义感知能力,弥合细粒度物理交互与高层推理之间的鸿沟,为可扩展、可解释、通用的具身智能铺平道路。

原文摘要 · Abstract (English)

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of embodied data fundamentally limit the alignment granularity between language and actions and exacerbate the challenge of long-horizon video generation--hindering generative models from achieving a "GPT moment" in the embodied domain. There is a naive observation: the diversity of embodied data far exceeds the relatively small space of possible primitive motions. Based on this insight, we propose a novel paradigm for world modeling--Primitive Embodied World Models (PEWM). By restricting video generation to fixed short horizons, our approach 1) enables fine-grained alignment between linguistic concepts and visual representations of robotic actions, 2) reduces learning complexity, 3) improves data efficiency in embodied data collection, and 4) decreases inference latency. By equipping with a modular Vision-Language Model (VLM) planner and a Start-Goal heatmap Guidance mechanism (SGG), PEWM further enables flexible closed-loop control and supports compositional generalization of primitive-level policies over extended, complex tasks. Our framework leverages the spatiotemporal vision priors in video models and the semantic awareness of VLMs to bridge the gap between fine-grained physical interaction and high-level reasoning, paving the way toward scalable, interpretable, and general-purpose embodied intelligence.

具身智能世界模型动作分解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。