让单视角3D高斯溅射可解释,自动分离几何与外观特征。
Interpretable Single-View 3D Gaussian Splatting using Unsupervised Hierarchical Disentangled Representation Learning
- 通过双分支结构分离3D几何与视觉外观,实现粗粒度解耦。
- 基于无监督解耦学习发现细粒度语义表示,保持高质量重建。
- 首个无监督可解释的单视图3D高斯溅射方法,适合需要可控生成的场景。
高斯溅射(GS)近期在3D重建中取得显著进展,实现快速渲染与高质量结果。然而,现有3DGS方法难以理解底层3D语义,限制了模型的可控性与可解释性。为此,我们提出可解释的单视图3DGS框架3DisGS,通过分层解耦表示学习(DRL)发现粗粒度与细粒度3D语义。模型采用双分支架构:点云初始化分支与三平面-高斯生成分支,实现几何与外观特征的粗粒度解耦。随后,利用基于DRL的编码器适配器进一步挖掘每种模态内的细粒度语义表示。据我们所知,这是首个实现无监督可解释3DGS的工作。实验表明,该模型在保持高质量与快速重建的同时实现了3D解耦。
原文摘要 · Abstract (English)
Gaussian Splatting (GS) has recently marked a significant advancement in 3D reconstruction, delivering both rapid rendering and high-quality results. However, existing 3DGS methods pose challenges in understanding underlying 3D semantics, which hinders model controllability and interpretability. To address it, we propose an interpretable single-view 3DGS framework, termed 3DisGS, to discover both coarse- and fine-grained 3D semantics via hierarchical disentangled representation learning (DRL). Specifically, the model employs a dual-branch architecture, consisting of a point cloud initialization branch and a triplane-Gaussian generation branch, to achieve coarse-grained disentanglement by separating 3D geometry and visual appearance features. Subsequently, fine-grained semantic representations within each modality are further discovered through DRL-based encoder-adapters. To our knowledge, this is the first work to achieve unsupervised interpretable 3DGS. Evaluations indicate that our model achieves 3D disentanglement while preserving high-quality and rapid reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。