arXiv:2512.06628cs.ROcs.CV2025-12被引 3

用认知模型生成长时序机器人操作视频,让虚拟数据更真实可用。

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

  • 分层设计:从任务规划到像素生成,跨层级协同建模。
  • 长序列合成效果领先,物理合理性通过强化学习提升。
  • 适合机器人政策训练与数据增强,无需人工轨迹标注。

可扩展的具身智能受限于多样且长时序的机器人操作数据稀缺。现有视频世界模型仅能合成简单动作的短片段,常依赖人工定义轨迹。为此,我们提出MIND-V,一种受认知科学启发的分层世界模型,用于合成物理合理且逻辑连贯的长时序机器人操作视频。MIND-V通过三个核心组件实现:语义推理中枢(SRH)利用预训练视觉语言模型进行任务规划;行为语义桥(BSB)将抽象指令转化为领域无关表示;运动视频生成器(MVG)实现条件视频渲染。采用分阶段视觉未来回滚策略提升长时序鲁棒性。为确保物理规律遵守,引入基于GRPO的强化学习后训练阶段,以新型物理预见一致性(PFC)奖励为指导。PFC利用V-JEPA2世界模型作为物理裁判,在潜在空间中惩罚不合理动态。实验表明,MIND-V在长时序仿真中表现领先,并显著提升策略学习性能,构建了完全自主、可扩展的具身数据合成框架。

原文摘要 · Abstract (English)

Scalable embodied intelligence is constrained by the scarcity of diverse, long-horizon robotic manipulation data. Existing video world models in this domain are limited to synthesizing short clips of simple actions and often rely on manually defined trajectories. To this end, we introduce MIND-V, a cognitive hierarchical world model designed to synthesize physically plausible and logically coherent videos of long-horizon robotic manipulation. Inspired by cognitive science, MIND-V bridges high-level reasoning with pixel-level synthesis through three core components: a Semantic Reasoning Hub (SRH) that leverages a pre-trained vision-language model for task planning; a Behavioral Semantic Bridge (BSB) that translates abstract instructions into domain-invariant representations; and a Motor Video Generator (MVG) for conditional video rendering. MIND-V employs Staged Visual Future Rollouts, a test-time optimization strategy to enhance long-horizon robustness. To enforce adherence to physical laws, we introduce a GRPO reinforcement learning post-training phase guided by a novel Physical Foresight Coherence (PFC) reward. PFC leverages the V-JEPA2 world model as a physics referee to penalize implausible dynamics in the latent feature space. Experiments confirm MIND-V's SOTA performance in long-horizon simulation and its significant value for policy learning, introducing a scalable and fully autonomous framework for embodied data synthesis.

机器人操作世界模型长时序生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。