arXiv:2511.18746cs.CVcs.AI2025-11

用基础动作生成4D内容,让语言与机器人动作精准对齐。

Any4D: Open-Prompt 4D Generation from Natural Language and Images

  • 限定短时视频生成,聚焦基础动作以提升对齐精度
  • 实现语言-动作细粒度匹配,支持复杂任务组合泛化
  • 适合需要高效、可解释的具身智能研究者

尽管基于视频生成的具身世界模型受到越来越多关注,但其对大规模具身交互数据的依赖仍是关键瓶颈。具身数据稀缺、采集困难且维度高,严重限制了语言与动作之间的对齐粒度,并加剧长时程视频生成的挑战,阻碍生成模型在具身领域实现类似GPT的突破。我们观察到:具身数据的多样性远超可能的基本动作空间。基于此,提出「基础具身世界模型」(PEWM),将视频生成限制在固定短时程内,该方法实现了:1)语言概念与机器人动作视觉表示的细粒度对齐;2)降低学习复杂度;3)提升具身数据收集的数据效率;4)减少推理延迟。通过引入模块化视觉-语言模型规划器与起始-目标热力图引导机制(SGG),PEWM进一步支持灵活闭环控制,并可在更长、复杂的任务中实现基础策略的组合泛化。框架利用视频模型中的时空视觉先验与视觉-语言模型的语义感知能力,弥合细粒度物理交互与高层推理之间的鸿沟,为可扩展、可解释、通用的具身智能铺平道路。

原文摘要 · Abstract (English)

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of embodied data fundamentally limit the alignment granularity between language and actions and exacerbate the challenge of long-horizon video generation--hindering generative models from achieving a \textit{"GPT moment"} in the embodied domain. There is a naive observation: \textit{the diversity of embodied data far exceeds the relatively small space of possible primitive motions}. Based on this insight, we propose \textbf{Primitive Embodied World Models} (PEWM), which restricts video generation to fixed shorter horizons, our approach \textit{1) enables} fine-grained alignment between linguistic concepts and visual representations of robotic actions, \textit{2) reduces} learning complexity, \textit{3) improves} data efficiency in embodied data collection, and \textit{4) decreases} inference latency. By equipping with a modular Vision-Language Model (VLM) planner and a Start-Goal heatmap Guidance mechanism (SGG), PEWM further enables flexible closed-loop control and supports compositional generalization of primitive-level policies over extended, complex tasks. Our framework leverages the spatiotemporal vision priors in video models and the semantic awareness of VLMs to bridge the gap between fine-grained physical interaction and high-level reasoning, paving the way toward scalable, interpretable, and general-purpose embodied intelligence.

具身智能视频生成视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。