arXiv:2605.01759cs.CV2026-05被引 1

通过跨样本语义传播,提升点云自监督学习的语义一致性。

PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning

论文配图:PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
图 1 · 摘自论文原文
  • 用状态空间模型串联批次样本,实现语义状态跨样本传递。
  • 在多个基准数据集上,性能与语义一致性均优于现有方法。
  • 适合需要稳定跨场景语义表达的3D视觉任务,如点云分类与分割。

场景级点云自监督学习(PC-SSL)在提升3D视觉模型泛化能力方面展现出潜力。然而,现有方法普遍采用样本独立建模范式,在不同场景间难以维持一致的语义表示,阻碍了统一且可迁移语义空间的构建。为此,我们提出基于跨样本语义传播(CSP)的PC-SSL框架:将批内样本序列化输入,由状态空间模型处理以实现语义状态传播,显式建模样本间的动态依赖关系,使网络在隐空间中建立跨样本语义一致性并达成全局语义对齐。由于序列化预训练依赖批级输入组织,我们进一步引入非对称语义保持蒸馏(SPD)机制用于微调,以实现语义迁移的结构对齐,并消除批依赖带来的不一致性。SPD通过异构输入机制和语义特征对齐约束,确保预训练语义的稳定迁移,使模型在单场景测试条件下仍能保持结构化语义一致性和鲁棒性。大量实验表明,该方法在多个基准数据集上持续优于当前最优方法,兼具更高性能与更强语义一致性。

原文摘要 · Abstract (English)

Scene-level point cloud self-supervised learning (PC-SSL) has demonstrated potential in enhancing the generalization capability of 3D vision models. Despite the advances in the field through existing methods, the sample-independent modeling paradigm still poses significant limitations in terms of maintaining consistent semantic representations across scenes. This challenge hinders the construction of a unified and transferable semantic space. To address this issue, we propose a PC-SSL framework based on cross-sample semantic propagation (CSP), in which samples within a batch are serialized into continuous input and processed by a state-space model to enable semantic state propagation. This mechanism explicitly models the dynamic dependencies across samples in the state space, allowing the network to establish cross-sample semantic consistency in the latent space and achieve global semantic alignment. Since serialization-based pretraining requires batch-level input organization, we further introduce an asymmetric semantic preservation distillation (SPD) during finetuning to achieve structural alignment of semantic transfer and eliminate inconsistencies caused by batch dependency. The proposed SPD ensures stable transfer of pretrained semantics through a heterogeneous input mechanism and a semantic feature alignment constraint. This enables the model to maintain structured semantic consistency and robustness under single-scene testing conditions. Extensive experiments on multiple benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods in both performance and semantic consistency.

点云学习自监督语义一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。