用双系统生成模型实现机器人实时规划与风险预警
MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
- 双扩散机制:低频生成场景,高频输出动作,提升效率
- 单步低分辨率特征控制,达11.3帧/秒,成功率63%(RL-Bench)
- 提前74%识别任务失败,适合高安全要求的实时机器人应用
视频生成模型(VGMs)已成为视觉-语言-动作(VLA)模型的核心,通过大规模预训练实现稳健的动力学建模。然而,现有方法未充分利用其分布建模能力进行未来状态预测。主要挑战在于:将生成过程融入特征学习在技术和概念上均不成熟;逐帧视频扩散计算效率低,难以满足机器人实时需求。为此,我们提出操纵梦境(MinD),一种用于实时、风险感知规划的双系统世界模型。MinD采用两个异步扩散过程:低频视觉生成器(LoDiff)预测未来场景,高频扩散策略(HiDiff)输出动作。核心洞察是:机器人策略无需完全去噪的图像,仅需单步去噪生成的低分辨率潜在表示即可。为连接早期预测与动作,我们引入DiffMatcher模块,采用新颖的协同训练策略同步两个扩散模型。MinD在RL-Bench上取得63%成功率,在真实世界Franka任务中达60%,运行速度为11.3 FPS,证明单步潜在特征可用于控制信号。此外,MinD可提前74%识别潜在任务失败,提供实时安全信号以支持监控与干预。本工作建立了基于生成世界模型的高效可靠机器人操作新范式。
原文摘要 · Abstract (English)
Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling capabilities for predicting future states. Two challenges hinder progress: integrating generative processes into feature learning is both technically and conceptually underdeveloped, and naive frame-by-frame video diffusion is computationally inefficient for real-time robotics. To address these, we propose Manipulate in Dream (MinD), a dual-system world model for real-time, risk-aware planning. MinD uses two asynchronous diffusion processes: a low-frequency visual generator (LoDiff) that predicts future scenes and a high-frequency diffusion policy (HiDiff) that outputs actions. Our key insight is that robotic policies do not require fully denoised frames but can rely on low-resolution latents generated in a single denoising step. To connect early predictions to actions, we introduce DiffMatcher, a video-action alignment module with a novel co-training strategy that synchronizes the two diffusion models. MinD achieves a 63% success rate on RL-Bench, 60% on real-world Franka tasks, and operates at 11.3 FPS, demonstrating the efficiency of single-step latent features for control signals. Furthermore, MinD identifies 74% of potential task failures in advance, providing real-time safety signals for monitoring and intervention. This work establishes a new paradigm for efficient and reliable robotic manipulation using generative world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。