arXiv:2601.02339cs.CV2026-01ICCV被引 2

让3D高斯模型同时懂语义和渲染,提升细节与效率

Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local Encoding

论文配图:Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local Encoding
图 1 · 摘自论文原文
  • 用各向异性切比雪夫描述符捕捉精细3D形状特征
  • 结合语义与形状信号动态调整高斯分布,提升渲染效率
  • 跨场景知识迁移加速收敛,适合新场景快速部署

近期工作尝试在3D高斯表示中加入语义特征向量,实现语义分割与图像渲染的同步。然而,这些方法通常将语义与渲染分支独立处理,仅依赖2D监督,忽视3D高斯几何结构。此外,现有自适应策略仅根据渲染梯度调整高斯点集,在细微或无纹理区域表现不足。本文提出一种联合增强框架,协同优化语义与渲染分支。首先,不同于传统点云形状编码,引入基于拉普拉斯-贝尔特拉米算子的各向异性3D高斯切比雪夫描述符,以捕捉精细3D形状细节,区分外观相似物体,降低对噪声2D引导的依赖。其次,不只依赖渲染梯度,通过局部语义与形状信号自适应调整高斯分配及球谐函数,实现资源的针对性分配,提升渲染效率。最后,设计跨场景知识迁移模块,持续更新学习到的形状模式,实现更快收敛与鲁棒表征,无需为每个新场景重新学习形状信息。在多个数据集上的实验表明,该方法在保持高帧率的同时,显著提升了分割准确率与渲染质量。

原文摘要 · Abstract (English)

Recent works propose extending 3DGS with semantic feature vectors for simultaneous semantic segmentation and image rendering. However, these methods often treat the semantic and rendering branches separately, relying solely on 2D supervision while ignoring the 3D Gaussian geometry. Moreover, current adaptive strategies adapt the Gaussian set depending solely on rendering gradients, which can be insufficient in subtle or textureless regions. In this work, we propose a joint enhancement framework for 3D semantic Gaussian modeling that synergizes both semantic and rendering branches. Firstly, unlike conventional point cloud shape encoding, we introduce an anisotropic 3D Gaussian Chebyshev descriptor using the Laplace-Beltrami operator to capture fine-grained 3D shape details, thereby distinguishing objects with similar appearances and reducing reliance on potentially noisy 2D guidance. In addition, without relying solely on rendering gradient, we adaptively adjust Gaussian allocation and spherical harmonics with local semantic and shape signals, enhancing rendering efficiency through selective resource allocation. Finally, we employ a cross-scene knowledge transfer module to continuously update learned shape patterns, enabling faster convergence and robust representations without relearning shape information from scratch for each new scene. Experiments on multiple datasets demonstrate improvements in segmentation accuracy and rendering quality while maintaining high rendering frame rates.

3D高斯语义分割渲染优化知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。