arXiv:2607.27749cs.CVcs.RO2026-07

仅凭闭合状态图像,就能重建可动物体的结构与关节。

Articulated Object Reconstruction from Rest-State Observation

论文配图:Articulated Object Reconstruction from Rest-State Observation
图 1 · 摘自论文原文
  • 用网格作为中间表示,融合视觉语言与分割结果
  • 通过扩散模型生成运动假设并验证几何一致性
  • 无需多姿态数据,适合数字孪生与交互建模

构建交互式数字孪生需恢复物体的3D几何结构及其运动关节。现有方法依赖多个姿态下的显式运动观测。本文提出一种仅需单个闭合状态的重构框架,在缺乏运动信息的病态条件下,利用几何、语义与运动先验进行补偿。采用显式网格作为中间表示,实现跨模型验证与融合,将视觉语言模型和分割模型的噪声输出整合为空间一致的部件结构。为估计关节参数,引入视频扩散模型生成运动假设,并通过几何一致性进行验证。本方法在部件分解与物理合理关节重建上表现优异,性能媲美依赖运动观测、生成式及模块化预训练模型的基线。

原文摘要 · Abstract (English)

Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.

物体重建关节估计扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。