arXiv:2412.18255cs.CV2024-12被引 10

用自适应方法修正视觉大模型噪声标签,提升3D语义分割精度

AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label Correction

  • 通过跨模态生成标签提供视觉大模型监督信号
  • 迭代更新噪声样本,显著降低户外复杂场景误差
  • 适合无标注数据的3D分割任务,尤其在开放环境表现优

近期,视觉基础模型(VFMs)在3D感知任务中展现出优异泛化能力。然而,在大规模室外数据集上,其性能受限于精确标注稀缺、外部环境变化带来的大量噪声以及未知物体众多等问题。本文提出一种全新的无标签学习方法——自适应标签修正(AdaCo)。AdaCo首先引入跨模态标签生成模块(CLGM),利用视觉基础模型的强大理解能力生成跨模态监督信号;随后,通过自适应噪声修正器(ANC)在训练过程中迭代更新并修正该监督信号中的噪声样本。此外,设计了自适应鲁棒损失函数(ARL),动态调节每个样本对噪声监督的敏感度,避免传统鲁棒损失导致的欠拟合问题。在两个室外基准数据集上的大量实验表明,AdaCo能有效克服无标签学习网络在3D语义分割中的性能瓶颈,显著提升分割精度。

原文摘要 · Abstract (English)

Recently, Visual Foundation Models (VFMs) have shown a remarkable generalization performance in 3D perception tasks. However, their effectiveness in large-scale outdoor datasets remains constrained by the scarcity of accurate supervision signals, the extensive noise caused by variable outdoor conditions, and the abundance of unknown objects. In this work, we propose a novel label-free learning method, Adaptive Label Correction (AdaCo), for 3D semantic segmentation. AdaCo first introduces the Cross-modal Label Generation Module (CLGM), providing cross-modal supervision with the formidable interpretive capabilities of the VFMs. Subsequently, AdaCo incorporates the Adaptive Noise Corrector (ANC), updating and adjusting the noisy samples within this supervision iteratively during training. Moreover, we develop an Adaptive Robust Loss (ARL) function to modulate each sample's sensitivity to noisy supervision, preventing potential underfitting issues associated with robust loss. Our proposed AdaCo can effectively mitigate the performance limitations of label-free learning networks in 3D semantic segmentation tasks. Extensive experiments on two outdoor benchmark datasets highlight the superior performance of our method.

3D分割视觉大模型噪声修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。