arXiv:2511.11700cs.CVeess.IV2025-11AAAI被引 1

无需预训练的点云语义分割模型,支持少样本与零样本场景。

EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance

论文配图:EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
图 1 · 摘自论文原文
  • 无预训练设计,通过原型增强注意力和双相对位置编码提取特征。
  • 引入语言引导原型嵌入,利用文本信息提升少样本性能并实现零样本推理。
  • 在S3DIS和ScanNet上分别超越当前最优方法5.68%和3.82%。

现有少样本3D点云语义分割方法通常采用两阶段学习流程,即先预训练再微调。尽管有效,但过度依赖预训练限制了模型灵活性与适应性。部分模型尝试避免预训练,却未能充分捕捉信息。此外,现有方法仅关注支持集中的视觉信息,忽视或未充分利用文本标注等辅助数据,导致性能受限且零样本能力弱。为此,本文提出一种新型无预训练网络EPSegFZ,包含三个核心组件:原型增强注册注意力(ProERA)模块与基于双相对位置编码(DRPE)的交叉注意力机制,实现无需预训练的特征提取与查询-原型精准匹配;语言引导原型嵌入(LGPE)模块,有效利用支持集中的文本信息,提升少样本表现并支持零样本推理。大量实验表明,该方法在S3DIS和ScanNet基准上分别领先当前最优方法5.68%和3.82%。

原文摘要 · Abstract (English)

Recent approaches for few-shot 3D point cloud semantic segmentation typically require a two-stage learning process, i.e., a pre-training stage followed by a few-shot training stage. While effective, these methods face overreliance on pre-training, which hinders model flexibility and adaptability. Some models tried to avoid pre-training yet failed to capture ample information. In addition, current approaches focus on visual information in the support set and neglect or do not fully exploit other useful data, such as textual annotations. This inadequate utilization of support information impairs the performance of the model and restricts its zero-shot ability. To address these limitations, we present a novel pre-training-free network, named Efficient Point Cloud Semantic Segmentation for Few- and Zero-shot scenarios. Our EPSegFZ incorporates three key components. A Prototype-Enhanced Registers Attention (ProERA) module and a Dual Relative Positional Encoding (DRPE)-based cross-attention mechanism for improved feature extraction and accurate query-prototype correspondence construction without pre-training. A Language-Guided Prototype Embedding (LGPE) module that effectively leverages textual information from the support set to improve few-shot performance and enable zero-shot inference. Extensive experiments show that our method outperforms the state-of-the-art method by 5.68% and 3.82% on the S3DIS and ScanNet benchmarks, respectively.

点云分割少样本学习语言引导零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。