arXiv:2512.01629cs.CVcs.RO2025-12被引 5

仅用一张图片生成可直接用于仿真的关节物体模型

SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge

  • 利用视觉语言模型提取粗略结构,再通过扩散Transformer生成零件与整体形状
  • 通过可微分正向运动学与渲染优化关节类型、轴线和位置,提升物理一致性
  • 支持多种物体类别,适合机器人抓取等下游仿真应用

关节式3D物体对具身AI、机器人和交互场景理解至关重要,但创建仿真可用资产仍需大量人工建模且依赖专家设计部件层级与运动结构。本文提出SPARK框架,仅凭单张RGB图像即可重建物理一致、具备运动结构的部件级关节物体。输入图像后,首先利用视觉语言模型(VLM)提取粗略的URDF参数并生成部件级参考图;随后将部件图像引导与推断出的结构图整合进生成式扩散Transformer,合成一致的部件与完整形状。为进一步优化URDF参数,引入可微分正向运动学与可微分渲染,在VLM生成的开状态监督下优化关节类型、轴线和原点。大量实验表明,SPARK可在多个类别上生成高质量、可直接用于仿真的关节资产,支持机器人操作与交互建模等下游应用。

原文摘要 · Abstract (English)

Articulated 3D objects are critical for embodied AI, robotics, and interactive scene understanding, yet creating simulation-ready assets remains labor-intensive and requires expert modeling of part hierarchies and motion structures. We introduce SPARK, a framework for reconstructing physically consistent, kinematic part-level articulated objects from a single RGB image. Given an input image, we first leverage VLMs to extract coarse URDF parameters and generate part-level reference images. We then integrate the part-image guidance and the inferred structure graph into a generative diffusion transformer to synthesize consistent part and complete shapes of articulated objects. To further refine the URDF parameters, we incorporate differentiable forward kinematics and differentiable rendering to optimize joint types, axes, and origins under VLM-generated open-state supervision. Extensive experiments show that SPARK produces high-quality, simulation-ready articulated assets across diverse categories, enabling downstream applications such as robotic manipulation and interaction modeling. Project page: https://heyumeng.com/SPARK/index.html.

3D重建关节物体扩散模型仿真生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。