用扩散模型+语义知识蒸馏,提升海洋目标分割精度
Marine Saliency Segmenter: Object-Focused Conditional Diffusion with Region-Level Semantic Knowledge Distillation
- 通过词级区域相似性匹配提取文本语义,指导条件特征学习
- 在MarineSaliencyDataset上达到86.7% mIoU,优于现有方法
- 适合需要精准分割海洋生物的科研与探测场景
海洋显著性分割(MSS)在各类基于视觉的海洋探测任务中至关重要。然而,现有方法因水下环境复杂,常出现目标定位偏差和边界模糊问题。尽管扩散模型在图像分割中表现优异,但其上下文语义利用仍有提升空间。为此,本文提出基于扩散模型的DiffMSS,通过区域级语义知识蒸馏增强显著对象的特征学习。具体而言,设计词级区域相似性匹配机制,从文本描述中识别出显著词汇,生成包含语义先验的条件特征;进一步引入专用一致性确定性采样策略,抑制对微细结构的过度自信误分割。大量实验表明,DiffMSS在定量与定性评估中均显著优于当前最优方法,在MarineSaliencyDataset上实现86.7%的mIoU。
原文摘要 · Abstract (English)
Marine Saliency Segmentation (MSS) plays a pivotal role in various vision-based marine exploration tasks. However, existing marine segmentation techniques face the dilemma of object mislocalization and imprecise boundaries due to the complex underwater environment. Meanwhile, despite the impressive performance of diffusion models in visual segmentation, there remains potential to further leverage contextual semantics to enhance feature learning of region-level salient objects, thereby improving segmentation outcomes. Building on this insight, we propose DiffMSS, a novel marine saliency segmenter based on the diffusion model, which utilizes semantic knowledge distillation to guide the segmentation of marine salient objects. Specifically, we design a region-word similarity matching mechanism to identify salient terms at the word level from the text descriptions. These high-level semantic features guide the conditional feature learning network in generating salient and accurate diffusion conditions with semantic knowledge distillation. To further refine the segmentation of fine-grained structures in unique marine organisms, we develop the dedicated consensus deterministic sampling to suppress overconfident missegmentations. Comprehensive experiments demonstrate the superior performance of DiffMSS over state-of-the-art methods in both quantitative and qualitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。