无需演示或强化学习,通过交互构建物体数字孪生实现精准操作
DexSim2Real$^{2}$: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation
- 通过自监督数据训练的显式交互网络预测动作
- 用多视角观测构建3D数字孪生,支持采样规划控制
- 适用于吸盘、夹爪和灵巧手,可拓展工具操作
关节类物体在日常生活中普遍存在。本文提出DexSim2Real²框架,实现目标导向的关节物体灵巧操作。核心是通过主动交互构建未见物体的显式世界模型,使基于采样的模型预测控制能在无示范或强化学习条件下规划达成不同目标的轨迹。系统首先利用自监督交互数据或人类操作视频训练的形态网络预测交互动作;执行后,提出一种基于3D AIGC的新建模流程,从多帧观测中在仿真中构建物体的数字孪生。针对灵巧手,采用特征抓取(eigengrasp)降低动作维度,提升轨迹搜索效率。实验验证了该框架在吸盘、两指夹爪及双灵巧手上的有效性,且显式世界模型具备泛化能力,支持工具操作等高级策略。
原文摘要 · Abstract (English)
Articulated objects are ubiquitous in daily life. In this paper, we present DexSim2Real$^{2}$, a novel framework for goal-conditioned articulated object manipulation. The core of our framework is constructing an explicit world model of unseen articulated objects through active interactions, which enables sampling-based model predictive control to plan trajectories achieving different goals without requiring demonstrations or RL. It first predicts an interaction using an affordance network trained on self-supervised interaction data or videos of human manipulation. After executing the interactions on the real robot to move the object parts, we propose a novel modeling pipeline based on 3D AIGC to build a digital twin of the object in simulation from multiple frames of observations. For dexterous hands, we utilize eigengrasp to reduce the action dimension, enabling more efficient trajectory searching. Experiments validate the framework's effectiveness for precise manipulation using a suction gripper, a two-finger gripper and two dexterous hand. The generalizability of the explicit world model also enables advanced manipulation strategies like manipulating with tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。