arXiv:2511.11563cs.CV2025-11SIGGRAPH被引 13

用稀疏图像一键重建带纹理和关节的3D可动物体,速度快质量高。

LARM: A Large Articulated-Object Reconstruction Model

  • 基于Transformer联合推断相机位姿与关节变化,端到端重建。
  • 在多个数据集上实现优于现有方法的视图合成与3D重建精度。
  • 适合需要快速高质量3D可动物体建模的研究与应用。

对具有真实几何、纹理和运动学特性的3D可动物体进行建模对于众多应用至关重要。然而,现有的基于优化的方法通常需要密集多视角输入和昂贵的实例级优化,限制了其可扩展性。近期的前馈方法虽更快,但常产生粗糙几何、缺乏纹理重建,并依赖脆弱复杂的多阶段流程。我们提出LARM,一个统一的前馈框架,通过联合恢复详细几何、真实纹理和准确关节结构,从稀疏视角图像重建3D可动物体。LARM将最近的静态物体新视角合成方法LVSM扩展至可动场景,利用基于Transformer的架构联合推理相机姿态与形变,实现可扩展且精确的新视角合成。此外,LARM生成深度图、部件掩码等辅助输出,便于显式提取3D网格和估计关节。该流程无需密集监督,支持跨多样物体类别的高保真重建。大量实验表明,LARM在新视角合成、状态重建及3D可动物体重建方面均超越现有最佳方法,生成高度贴合输入图像的高质量网格。项目页面:https://sylviayuan-sy.github.io/larm-site/

原文摘要 · Abstract (English)

Modeling 3D articulated objects with realistic geometry, textures, and kinematics is essential for a wide range of applications. However, existing optimization-based reconstruction methods often require dense multi-view inputs and expensive per-instance optimization, limiting their scalability. Recent feedforward approaches offer faster alternatives but frequently produce coarse geometry, lack texture reconstruction, and rely on brittle, complex multi-stage pipelines. We introduce LARM, a unified feedforward framework that reconstructs 3D articulated objects from sparse-view images by jointly recovering detailed geometry, realistic textures, and accurate joint structures. LARM extends LVSM a recent novel view synthesis (NVS) approach for static 3D objects into the articulated setting by jointly reasoning over camera pose and articulation variation using a transformer-based architecture, enabling scalable and accurate novel view synthesis. In addition, LARM generates auxiliary outputs such as depth maps and part masks to facilitate explicit 3D mesh extraction and joint estimation. Our pipeline eliminates the need for dense supervision and supports high-fidelity reconstruction across diverse object categories. Extensive experiments demonstrate that LARM outperforms state-of-the-art methods in both novel view and state synthesis as well as 3D articulated object reconstruction, generating high-quality meshes that closely adhere to the input images. project page: https://sylviayuan-sy.github.io/larm-site/

3D重建可动物体新视角合成视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。