arXiv:2603.13825cs.RO2026-03

用数字孪生实现零样本开放世界物体操作,无需示范即可通用

Building Explicit World Model for Zero-Shot Open-World Object Manipulation

  • 构建物理可信的环境数字孪生,替代依赖大量示范
  • 零样本完成多类物体和任务的操控,真实世界部署可靠
  • 适合研究开放世界机器人操控与具身智能的学者

开放世界物体操作仍是机器人领域的基础挑战。尽管视觉-语言-动作(VLA)模型表现良好,但其严重依赖大规模机器人动作示范,收集成本高且易导致分布外泛化能力差。本文提出一种基于显式世界模型的开放世界操控框架,通过构建环境的物理可信数字孪生,实现零样本泛化。该框架融合开放集感知、数字孪生重建以及交互策略采样与评估。借助物理仿真器,高效探索并验证操控策略,最终可靠部署至真实世界。实验表明,该框架无需任何特定任务的动作示范,即可完成多种开放集操控任务,在任务与物体层面均展现出强大零样本泛化能力。

原文摘要 · Abstract (English)

Open-world object manipulation remains a fundamental challenge in robotics. While Vision-Language-Action (VLA) models have demonstrated promising results, they rely heavily on large-scale robot action demonstrations, which are costly to collect and can hinder out-of-distribution generalization. In this paper, we propose an explicit-world-model-based framework for open-world manipulation that achieves zero-shot generalization by constructing a physically grounded digital twin of the environment. The framework integrates open-set perception, digital-twin reconstruction, sampling and evaluation of interaction strategies. By constructing a digital twin of the environment, our approach efficiently explores and evaluates manipulation strategies in physic-enabled simulator and reliably deploys the chosen strategy to the real world. Experimentally, the proposed framework is able to perform multiple open-set manipulation tasks without any task-specific action demonstrations, proving strong zero-shot generalization on both the task and object levels. Project Page: https://bojack-bj.github.io/projects/thesis/

机器人操控数字孪生零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。