针对声呐图像标注极少时的语义分割难题,提出多教师协作框架提升模型性能。
CTFS : Collaborative Teacher Framework for Forward-Looking Sonar Image Semantic Segmentation with Extremely Limited Labels
- 设计多教师协作机制,融合通用与声呐特化教师指导学生学习。
- 在仅2%标注数据下,相比现有方法提升5.08% mIoU,显著改善分割效果。
- 通过跨教师一致性评估减少噪声伪标签干扰,适合小样本声呐图像分析。
作为重要的水下感知技术,前视声呐具有独特的成像特性。声呐图像常受严重斑点噪声、低纹理对比度、声学阴影和几何畸变影响,导致传统师生框架在极少量标注数据下难以取得满意性能。为此,我们提出一种面向前视声呐图像的协同教师语义分割框架。该框架引入由一个通用教师和多个声呐特化教师组成的多教师协作机制,采用交替引导策略,使学生模型既能学习通用语义表征,又能捕捉声呐图像的独特特征,实现更全面、鲁棒的特征建模。针对声呐图像带来的伪标签噪声问题,我们进一步设计了跨教师可靠性评估机制,通过多视角、多教师预测的一致性与稳定性动态量化伪标签可靠性,有效缓解噪声伪标签的负面影响。值得注意的是,在FLSMD数据集上,当仅有2%数据标注时,本方法相较其他先进方法实现了5.08%的mIoU提升。
原文摘要 · Abstract (English)
As one of the most important underwater sensing technologies, forward-looking sonar exhibits unique imaging characteristics. Sonar images are often affected by severe speckle noise, low texture contrast, acoustic shadows, and geometric distortions. These factors make it difficult for traditional teacher-student frameworks to achieve satisfactory performance in sonar semantic segmentation tasks under extremely limited labeled data conditions. To address this issue, we propose a Collaborative Teacher Semantic Segmentation Framework for forward-looking sonar images. This framework introduces a multi-teacher collaborative mechanism composed of one general teacher and multiple sonar-specific teachers. By adopting a multi-teacher alternating guidance strategy, the student model can learn general semantic representations while simultaneously capturing the unique characteristics of sonar images, thereby achieving more comprehensive and robust feature modeling. Considering the challenges of sonar images, which can lead teachers to generate a large number of noisy pseudo-labels, we further design a cross-teacher reliability assessment mechanism. This mechanism dynamically quantifies the reliability of pseudo-labels by evaluating the consistency and stability of predictions across multiple views and multiple teachers, thereby mitigating the negative impact caused by noisy pseudo-labels. Notably, on the FLSMD dataset, when only 2% of the data is labeled, our method achieves a 5.08% improvement in mIoU compared to other state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。