arXiv:2601.07823eess.SYcs.RO2026-01被引 15

视频生成模型助力机器人模拟真实世界,提升仿真精度与智能决策能力。

Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions

论文配图:Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions
图 1 · 摘自论文原文
  • 用多模态输入生成高保真视频,替代传统物理引擎简化假设
  • 支持模仿学习中的动作预测与强化学习的动力学建模,提升策略性能
  • 适合关注机器人仿真、视觉规划与安全可控生成的研究者

视频生成模型作为高保真物理世界建模工具,能基于多模态用户输入合成高质量视频,捕捉代理与环境间的精细交互。其能力克服了物理模拟器长期存在的瓶颈,如对可变形体的逼真模拟无需过度简化假设。视频模型可作为基础世界模型,以细粒度方式表达世界动态,超越语言抽象对复杂物理交互的描述局限。本文综述视频模型在机器人领域的应用,涵盖模仿学习中的低成本数据生成与动作预测、强化学习中的动力学与奖励建模、视觉规划及策略评估。同时指出其可信集成面临挑战:指令遵循差、出现违反物理规律的幻觉、生成不安全内容,以及高昂的数据整理、训练和推理成本。提出未来研究方向以推动其在安全关键场景中的广泛应用。

原文摘要 · Abstract (English)

Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents and their environments conditioned on multi-modal user inputs. Their impressive capabilities address many of the long-standing challenges faced by physics-based simulators, driving broad adoption in many problem domains, e.g., robotics. For example, video models enable photorealistic, physically consistent deformable-body simulation without making prohibitive simplifying assumptions, which is a major bottleneck in physics-based simulation. Moreover, video models can serve as foundation world models that capture the dynamics of the world in a fine-grained and expressive way. They thus overcome the limited expressiveness of language-only abstractions in describing intricate physical interactions. In this survey, we provide a review of video models and their applications as embodied world models in robotics, encompassing cost-effective data generation and action prediction in imitation learning, dynamics and rewards modeling in reinforcement learning, visual planning, and policy evaluation. Further, we highlight important challenges hindering the trustworthy integration of video models in robotics, which include poor instruction following, hallucinations such as violations of physics, and unsafe content generation, in addition to fundamental limitations such as significant data curation, training, and inference costs. We present potential future directions to address these open research challenges to motivate research and ultimately facilitate broader applications, especially in safety-critical settings.

视频生成机器人世界模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。