arXiv:2603.14010cs.RO2026-03

从一张图生成可运行的机械结构模型,无需分步处理

URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets

  • 端到端扩散模型直接生成带关节参数的完整URDF
  • 在多个基准上几何重建和物理可执行性均超越旧方法
  • 适合需要快速构建仿真资产的机器人研究者

关节物体是机器人、物理模拟和交互式虚拟环境的基础。然而,仅从视觉观测中恢复它们极具挑战性,因为图像仅提供关于部件几何与运动结构的局部且模糊的信息。现有方法通常依赖多阶段流程、从资产库检索或显式部件分割。本文提出URDF-Anything+,一种端到端自回归扩散框架,可直接从单张RGB图像生成可运行的URDF模型。该模型在结构化潜在空间中运作,联合建模部件几何与关节结构。具体而言,模型按顺序预测每个关节部件及其关联的关节参数,同时通过终止标记动态决定部件数量。这一设计使模型可直接生成完全可执行的URDF,无需外部检索或后处理。在大规模关节物体基准上的实验表明,URDF-Anything+在几何重建质量、关节参数估计与物理可执行性方面均优于现有方法,且显著优于多阶段方法的效率。此外,生成的URDF可作为忠实数字孪生,实现纯仿真训练的抓取策略零样本迁移。

原文摘要 · Abstract (English)

Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual observations is inherently challenging, as images provide only partial and ambiguous cues about both part geometry and their underlying kinematic structure. Existing approaches typically rely on multi-stage pipelines, retrieval from asset libraries, or explicit part segmentation. We present URDF-Anything+, an end-to-end autoregressive diffusion framework that generates simulation-ready URDF models directly from a single RGB image. Conditioned on visual observations and object geometry, URDF-Anything+ operates in a structured latent space and jointly models part geometry and articulation in a unified generation process. Specifically, the model sequentially predicts each articulated part together with its associated joint parameters, while a termination token dynamically determines the number of parts. This design enables direct generation of fully executable URDFs without external retrieval or post-processing stages. Experiments on large-scale articulated object benchmarks demonstrate that URDF-Anything+ outperforms prior methods in geometric reconstruction quality, joint parameter estimation, and physical executability, while being substantially more efficient than existing multi-stage approaches. Furthermore, the generated URDFs serve as faithful digital twins, enabling the zero-shot transfer of manipulation policies trained purely in simulation.

生成模型机器人仿真数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。