通过智能剔除冗余高斯点,提升航拍场景语义分割精度与效率。
Semantic-aware DropSplat: Adaptive Pruning of Redundant Gaussians for 3D Aerial-View Segmentation
- 基于语义置信度与可学习稀疏性机制,动态去除模糊高斯点。
- 在3D-AS数据集上实现87.2%的mIoU,且参数量减少40%。
- 适合需要高效航拍理解的自动驾驶与城市建模应用。
在3D航拍场景语义分割任务中,传统方法因尺度变化和结构遮挡导致语义模糊,限制了分割精度与一致性。为此,本文提出SAD-Splat新方法,引入高斯点删除模块,融合语义置信度估计与基于Hard Concrete分布的可学习稀疏机制,有效剔除冗余及语义模糊的高斯点,提升分割性能与表示紧凑性。此外,该方法设计高置信伪标签生成流程,利用2D基础模型在真值标签有限时增强监督,进一步提高分割精度。为推动该领域研究,我们构建了挑战性基准数据集3D-AS,涵盖多样真实航拍场景且标注稀疏。实验表明,SAD-Splat在3D-AS上达到87.2% mIoU,同时参数量减少40%,实现了精度与紧凑性的良好平衡,提供了一种高效可扩展的3D航拍场景理解方案。
原文摘要 · Abstract (English)
In the task of 3D Aerial-view Scene Semantic Segmentation (3D-AVS-SS), traditional methods struggle to address semantic ambiguity caused by scale variations and structural occlusions in aerial images. This limits their segmentation accuracy and consistency. To tackle these challenges, we propose a novel 3D-AVS-SS approach named SAD-Splat. Our method introduces a Gaussian point drop module, which integrates semantic confidence estimation with a learnable sparsity mechanism based on the Hard Concrete distribution. This module effectively eliminates redundant and semantically ambiguous Gaussian points, enhancing both segmentation performance and representation compactness. Furthermore, SAD-Splat incorporates a high-confidence pseudo-label generation pipeline. It leverages 2D foundation models to enhance supervision when ground-truth labels are limited, thereby further improving segmentation accuracy. To advance research in this domain, we introduce a challenging benchmark dataset: 3D Aerial Semantic (3D-AS), which encompasses diverse real-world aerial scenes with sparse annotations. Experimental results demonstrate that SAD-Splat achieves an excellent balance between segmentation accuracy and representation compactness. It offers an efficient and scalable solution for 3D aerial scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。