arXiv:2608.07074cs.RO2026-08

用分层多模型表示,让机器人在低内存下高效建图。

M2-SMap: Memory-Efficient Semantic Mapping with Hierarchical Multi-Model Representation

论文配图:M2-SMap: Memory-Efficient Semantic Mapping with Hierarchical Multi-Model Representation
图 1 · 摘自论文原文
  • 分层分解点云为紧凑高斯体,提升结构表达灵活性。
  • 实测帧率超29.37Hz,原生基线减少18.7%的图元数量。
  • 有效消除物体粘连,实现语义一致的轻量化建图。

密集点云地图在资源受限机器人上部署困难,因内存随场景规模迅速增长。尽管紧凑单模型表示降低内存开销,但其固定几何表达能力不足以应对结构多样的环境。现有多模型方法虽提升表达灵活性,但特征提取与模型选择常依赖局部几何,易导致过拟合和物体粘连。本文提出M2-SMap,基于分层多模型表示的高效语义建图框架。首先,分层几何分解将RGB-D点云划分为紧凑高斯组件;其次,投影引导的语义标注机制为每个组件分配实例身份;随后,将标注信息融入面向对象的高斯融合策略。此外,多尺度特征提取分离大平面区域、语义物体与复杂残差结构,分别以有界平面、物体级超二次曲面和高斯混合模型(GMM)原型表示。在三个RGB-D序列上的实验表明,M2-SMap实现不低于29.37 Hz的实时运行,原始基线平均减少18.7%的图元数量;同时,每帧物体间粘连测量数从2.808降至0,实现高效且语义一致的场景表示。

原文摘要 · Abstract (English)

Dense point cloud maps, as a typically used mapping representation, are difficult to deploy on resource-constrained robots because their memory consumption grows rapidly with scene scale. Although compact single-model representations reduce memory cost, their fixed geometric expressiveness is insufficient for structurally diverse environments. Existing multi-model methods improve representational flexibility, yet their feature extraction and model selection are often dominated by local geometry, which can cause overfitting and adhesion between objects. To address these issues, this paper presents M2-SMap, a memory-efficient semantic mapping framework based on hierarchical multi-model representation. First, a hierarchical geometric decomposition partitions RGB-D point clouds into compact Gaussian components. Then, a projection-guided semantic annotation mechanism assigns instance identities to each component. Subsequently, these annotations are incorporated into an object-aware Gaussian fusion strategy. Furthermore, a multi-scale feature extraction strategy separates large planar regions, semantic objects, and complex residual structures, which are respectively represented by bounded planes, object-level superquadrics, and GMM primitives. Experiments on three RGB-D sequences show that M2-SMap runs in real time at no less than 29.37 Hz while achieving the lowest primitive count, with an average reduction of 18.7% over the best baseline. It also reduces the mean per-frame number of measured inter-object adhesion cases from 2.808 to 0, demonstrating efficient and semantically consistent scene representation.

语义建图轻量化高斯表示多模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。