用文本引导融合语义信息,提升3D高斯点云重建的细节质量
TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
- 通过文本引导的注意力机制融合深度、语义与多视角特征
- 在多个基准数据集上优于现有方法,显著提升重建细节精度
- 适合需要细粒度语义理解的3D重建任务,如复杂场景建模
通用高斯点云渲染近年来通过前馈模型实现了从稀疏视图中鲁棒重建3D场景,具备出色的跨场景泛化能力。然而,多数方法侧重几何一致性,忽视了文本驱动的语义引导对复杂场景精细结构重建的关键作用。为此,我们提出TextSplat——首个文本驱动的通用高斯点云框架。该框架通过文本引导融合多种语义线索,学习跨模态鲁棒特征表示,增强几何与语义信息对齐,生成高保真3D重建。具体地,采用三个并行模块分别获取互补表征:扩散先验深度估计器用于精确深度,语义感知分割网络提供细粒度语义,多视角交互网络优化跨视角特征。随后,在文本引导的语义融合模块中,通过文本引导的注意力机制聚合这些表征,生成富含语义细节的3D高斯参数。在多个基准数据集上的实验表明,该框架在多项评估指标上均优于现有方法,验证了其有效性。代码将公开。
原文摘要 · Abstract (English)
Recent advancements in Generalizable Gaussian Splatting have enabled robust 3D reconstruction from sparse input views by utilizing feed-forward Gaussian Splatting models, achieving superior cross-scene generalization. However, while many methods focus on geometric consistency, they often neglect the potential of text-driven guidance to enhance semantic understanding, which is crucial for accurately reconstructing fine-grained details in complex scenes. To address this limitation, we propose TextSplat--the first text-driven Generalizable Gaussian Splatting framework. By employing a text-guided fusion of diverse semantic cues, our framework learns robust cross-modal feature representations that improve the alignment of geometric and semantic information, producing high-fidelity 3D reconstructions. Specifically, our framework employs three parallel modules to obtain complementary representations: the Diffusion Prior Depth Estimator for accurate depth information, the Semantic Aware Segmentation Network for detailed semantic information, and the Multi-View Interaction Network for refined cross-view features. Then, in the Text-Guided Semantic Fusion Module, these representations are integrated via the text-guided and attention-based feature aggregation mechanism, resulting in enhanced 3D Gaussian parameters enriched with detailed semantic cues. Experimental results on various benchmark datasets demonstrate improved performance compared to existing methods across multiple evaluation metrics, validating the effectiveness of our framework. The code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。