arXiv:2511.17092cs.CV2025-11

仅用单状态稀疏视角图像,实现关节物体高精度3D重建。

SPAGS: Sparse-View Articulated Object Reconstruction from Single State via Planar Gaussian Splatting

  • 用平面高斯点云替代传统高斯,提升几何精度。
  • 在合成与真实数据集上,部分级表面重建效果优于现有方法。
  • 结合视觉语言模型,实现开放词汇部件分割与参数联合估计。

关节物体广泛存在于日常环境中,其三维重建在多个领域具有重要意义。然而,现有方法通常需要昂贵的多阶段、多视角观测输入。为此,我们提出一种无需类别先验的关节物体重建框架,基于平面高斯点阵,仅需单状态下的稀疏视角RGB图像。首先,引入高斯信息场从候选相机位姿中感知最优稀疏视点;为确保精确几何保真度,将传统3D高斯约束为平面基元,以提升法向量与深度估计精度。平面高斯点通过粗到精的方式优化,并受深度平滑性与少样本扩散先验正则化。此外,借助视觉语言模型(VLM)通过视觉提示实现开放词汇部件分割与联合参数估计。在合成与真实数据集上的大量实验表明,本方法显著优于现有基线,在部件级表面重建保真度方面表现优异。

原文摘要 · Abstract (English)

Articulated objects are ubiquitous in daily environments, and their 3D reconstruction holds great significance across various fields. However, existing articulated object reconstruction methods typically require costly inputs such as multi-stage and multi-view observations. To address the limitations, we propose a category-agnostic articulated object reconstruction framework via planar Gaussian Splatting, which only uses sparse-view RGB images from a single state. Specifically, we first introduce a Gaussian information field to perceive the optimal sparse viewpoints from candidate camera poses. To ensure precise geometric fidelity, we constrain traditional 3D Gaussians into planar primitives, facilitating accurate normal and depth estimation. The planar Gaussians are then optimized in a coarse-to-fine manner, regularized by depth smoothness and few-shot diffusion priors. Furthermore, we leverage a Vision-Language Model (VLM) via visual prompting to achieve open-vocabulary part segmentation and joint parameter estimation. Extensive experiments on both synthetic and real-world datasets demonstrate that our approach significantly outperforms existing baselines, achieving superior part-level surface reconstruction fidelity.

3D重建平面高斯关节物体少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。