arXiv:2511.17904cs.CVcs.RO2025-11

用600万参数实现多模态3D场景的高效统一建模

CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation

  • 构建体素锚点结构连接语义与几何
  • 仅需600万参数即达顶尖性能
  • 适合追求轻量化多模态3D建模的研究者

基于高斯点阵的3D场景表示近年呈现两大趋势:侧重语义理解但缺乏显式几何建模的方法,以及捕捉空间结构却语义抽象不足的方法。为弥合这一差距,我们提出CUS-GS——一种紧凑统一的结构化高斯点阵框架,将多模态语义特征与结构化3D几何相联接。具体而言,设计体素化锚点结构构建空间骨架,并从基础模型(如CLIP、DINOv2、SEEM)中提取多模态语义特征。引入多模态潜在特征分配机制,统一跨异构特征空间的外观、几何与语义表示,确保多模型间的一致性。此外,提出特征感知显著性评估策略,动态引导锚点的生长与裁剪,有效剔除冗余或无效锚点,同时保持语义完整性。大量实验表明,CUS-GS在仅使用600万参数的情况下,性能媲美当前最优方法,较最接近的对手(3500万参数)低一个数量级,充分体现了该框架在性能与模型效率间的优异平衡。

原文摘要 · Abstract (English)

Recent advances in Gaussian Splatting based 3D scene representation have shown two major trends: semantics-oriented approaches that focus on high-level understanding but lack explicit 3D geometry modeling, and structure-oriented approaches that capture spatial structures yet provide limited semantic abstraction. To bridge this gap, we present CUS-GS, a compact unified structured Gaussian Splatting representation, which connects multimodal semantic features with structured 3D geometry. Specifically, we design a voxelized anchor structure that constructs a spatial scaffold, while extracting multimodal semantic features from a set of foundation models (e.g., CLIP, DINOv2, SEEM). Moreover, we introduce a multimodal latent feature allocation mechanism to unify appearance, geometry, and semantics across heterogeneous feature spaces, ensuring a consistent representation across multiple foundation models. Finally, we propose a feature-aware significance evaluation strategy to dynamically guide anchor growing and pruning, effectively removing redundant or invalid anchors while maintaining semantic integrity. Extensive experiments show that CUS-GS achieves competitive performance compared to state-of-the-art methods using as few as 6M parameters - an order of magnitude smaller than the closest rival at 35M - highlighting the excellent trade off between performance and model efficiency of the proposed framework.

3D建模多模态轻量化高斯点阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。