arXiv:2607.04546cs.ROcs.AI2026-07被引 2

用分割掩码做桥梁,让仿真训练的机械手模型更贴近真实操作。

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models

论文配图:Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
图 1 · 摘自论文原文
  • 分两阶段建模:先预测分割掩码,再生成真实图像。
  • 用50小时仿真数据预训练,仅需2.5小时实操微调,实现23自由度精确控制。
  • 适合需要精细动作控制的机器人研究者,提升仿真到现实的迁移效果。

动作条件世界模型使机器人能在不进行额外物理交互的情况下预测未来行为结果,支持策略评估、规划和数据增强。我们提出Mask2Real-WM,一种用于灵巧操作的两阶段动作条件世界模型,将像素预测解耦为动力学模型和渲染模型。动力学模型基于历史分割掩码和23自由度动作序列,预测未来分割掩码;渲染模型则利用ControlNet增强的Stable Video Diffusion主干网络,将预测掩码映射为逼真RGB图像。由于分割空间中的仿真到现实差距较小,动力学模型可借助超过50小时的合成仿真数据进行大规模预训练,随后在不足2.5小时的真实示范数据上微调。在灵巧抓取与放置基准测试中,掩码条件和仿真预训练均对全部23个自由度的每自由度动作可控性至关重要。相比之下,单体基线模型虽能捕捉整体手部与末端执行器轨迹,但无法可靠反映细粒度的单关节动作影响。

原文摘要 · Abstract (English)

Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy evaluation, planning, and data augmentation. We present Mask2Real-WM, a two-stage action-conditioned world model for dexterous manipulation that decouples pixel prediction into a dynamics model and a rendering model. The dynamics model predicts future segmentation masks from past masks and 23-DoF action sequences. The rendering model maps the predicted masks to photorealistic RGB using a ControlNet-augmented Stable Video Diffusion backbone. The smaller sim-to-real gap in segmentation space enables the dynamics model to benefit from large-scale pretraining on over 50 h of synthetic simulation data, followed by fine-tuning on fewer than 2.5 h of real demonstrations. Experiments on a dexterous pick-and-place benchmark show that mask conditioning and simulation pretraining are both required for per-DoF action controllability across all 23 degrees of freedom. In contrast, monolithic baselines capture broad hand and end-effector trajectories but do not reliably reflect fine-grained, per-joint action effects.

世界模型灵巧操作仿真到现实分割掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。