提升3D点云未知类别识别能力,让模型能发现没见过的新物体
Improving Open-Set Semantic Segmentation in 3D Point Clouds by Conditional Channel Capacity Maximization: Preliminary Results
- 将分割过程建模为条件马尔可夫链,设计新正则项增强特征区分度
- 在多个数据集上显著提升对未见类别的检测能力,验证了方法有效性
- 适合需要动态适应新物体的自动驾驶、机器人等场景
点云语义分割支撑众多关键应用。尽管近期深度网络和大规模数据集推动了闭集性能的飞跃,但这些模型难以识别或正确分割训练类别之外的物体。这一差距催生了开放集语义分割(O3S),要求模型既能准确标注已知类别,又能检测未知类别。本文提出一种即插即用的O3S框架。通过将分割流程建模为条件马尔可夫链,推导出一种名为条件通道容量最大化(3CM)的新正则项,该正则项在每类条件下最大化特征与预测间的互信息。将其融入标准损失函数后,3CM促使编码器保留更丰富、与标签相关的特征,从而增强网络对先前未见类别的区分与分割能力。实验结果表明该方法在检测未见物体方面具有显著效果。同时,论文还展望了未来在动态开放世界自适应及高效信息论估计方面的方向。
原文摘要 · Abstract (English)
Point-cloud semantic segmentation underpins a wide range of critical applications. Although recent deep architectures and large-scale datasets have driven impressive closed-set performance, these models struggle to recognize or properly segment objects outside their training classes. This gap has sparked interest in Open-Set Semantic Segmentation (O3S), where models must both correctly label known categories and detect novel, unseen classes. In this paper, we propose a plug and play framework for O3S. By modeling the segmentation pipeline as a conditional Markov chain, we derive a novel regularizer term dubbed Conditional Channel Capacity Maximization (3CM), that maximizes the mutual information between features and predictions conditioned on each class. When incorporated into standard loss functions, 3CM encourages the encoder to retain richer, label-dependent features, thereby enhancing the network's ability to distinguish and segment previously unseen categories. Experimental results demonstrate effectiveness of proposed method on detecting unseen objects. We further outline future directions for dynamic open-world adaptation and efficient information-theoretic estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。