用多模态世界模型提升机器人抓握的泛化与鲁棒性
WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

- 基于触觉、深度等多源感知,学习动作相关的紧凑动态状态
- 预训练模型支持49种物体的下游任务,提升未见物体操作能力
- 通过去噪深度重建,实现真实机器人部署的稳定抓握
泛化且鲁棒的灵巧手部操作需要策略从部分和噪声观测中推断物体位姿、几何、接触及潜在滑移。尽管近期触觉与视觉触觉强化学习方法在控制环境下表现优异,但在位姿变化、力扰动和物体差异下鲁棒性下降。我们提出WM-Craftnet,一种由世界模型驱动的框架,从本体感知、深度、触觉和动作中学习动作条件下的紧凑潜在动态,通过多模态重构和奖励预测进行监督。不同于将世界模型用于潜在想象或策略优化,WM-Craftnet将学习到的世界联觉模型(WSM)作为非对称演员-评论家策略的递归任务上下文。重要的是,WSM被训练为从噪声深度输入中重构干净深度目标,为真实机器人部署提供去噪几何状态。消融实验表明,预测性世界建模、干净深度监督和触觉接触线索共同塑造了学习状态。在九种z轴物体上预训练的WSM可作为49种物体下游策略学习的通用先验,显著提升多物体旋转性能,定量与定性证据显示其具备未见物体操作、扰动恢复及仿真到真实迁移能力。
原文摘要 · Abstract (English)
Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, contact, and potential slip from partial and noisy observations. Although recent tactile and visuotactile RL methods achieve strong in-hand rotation in controlled settings, their robustness often degrades under pose shifts, force disturbances, and object variation. We propose WM-Craftnet, a world-model-conditioned framework that learns compact action-conditioned latent dynamics from proprioception, depth, tactile sensing, and actions, supervised by multimodal reconstruction and reward prediction. Rather than using the world model for latent imagination or policy optimization, WM-Craftnet uses the learned World Synesthesia Model (WSM) as recurrent task context for an asymmetric actor--critic policy. Importantly, WSM is trained to reconstruct clean depth targets from noisy depth inputs, providing a denoised geometric state for real-robot deployment. Ablations over recurrent baselines, auxiliary heads, tactile masking, and WSM modality heads show that predictive world modeling, clean-depth supervision, and tactile contact cues all shape the learned state. A WSM pretrained on nine \(z\)-axis objects serves as a reusable prior for \(49\)-object downstream policy learning. This context improves multi-object rotation, with quantitative and qualitative evidence for unseen-object, perturbation-recovery, and sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。