arXiv:2603.19231cs.CV2026-03被引 4

单图重建可动3D物体,通过逐步推理结构提升精度与效率

MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction

  • 分步推理:从图像逐步构建标准几何、部件结构和运动感知表征
  • 在PartNet-Mobility上达到最优重建精度与最快推理速度
  • 无需外部运动模板,适用于机器人操作与复杂场景重建

从单张图像重建可动3D物体需同时推断几何形状、部件结构与运动参数,核心难点在于运动线索与结构之间的纠缠,导致直接回归关节参数不稳定。现有方法依赖多视角监督、基于检索的组装或辅助视频生成,常牺牲可扩展性或效率。本文提出MonoArt,一个基于渐进式结构推理的统一框架。不直接从图像特征预测关节,而是通过单一架构逐步将视觉观察转化为标准几何、结构化部件表示和运动感知嵌入。该结构化推理过程实现稳定且可解释的关节推断,无需外部运动模板或多阶段流程。在PartNet-Mobility上的大量实验表明,MonoArt在重建精度与推理速度上均达当前最优。框架还成功推广至机器人操作与可动场景重建任务。

原文摘要 · Abstract (English)

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and object structure, which makes direct articulation regression unstable. Existing methods address this challenge through multi-view supervision, retrieval-based assembly, or auxiliary video generation, often sacrificing scalability or efficiency. We present MonoArt, a unified framework grounded in progressive structural reasoning. Rather than predicting articulation directly from image features, MonoArt progressively transforms visual observations into canonical geometry, structured part representations, and motion-aware embeddings within a single architecture. This structured reasoning process enables stable and interpretable articulation inference without external motion templates or multi-stage pipelines. Extensive experiments on PartNet-Mobility demonstrate that OM achieves state-of-the-art performance in both reconstruction accuracy and inference speed. The framework further generalizes to robotic manipulation and articulated scene reconstruction.

3D重建单图重建结构推理可动物体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。