arXiv:2503.16254cs.CV2025-03

无需训练的多模态交互分割,用深度图提升精度与稳定性

M2N2V2: Multi-Modal Unsupervised and Training-free Interactive Segmentation

  • 结合深度图与注意力图构建新型马尔可夫映射,实现无监督交互分割
  • 提出自适应评分函数,减少点击过程中分割区域大小波动,提升mIoU
  • 在DAVIS和HQSeg44K上性能接近有监督方法,适合无需标注数据的场景

我们提出马尔可夫映射最近邻第二版(M2N2V2),一种新颖且简单但高效的无监督、无需训练的点提示交互分割方法。借鉴监督式多模态方法趋势,我们引入深度作为额外模态,构建新的深度引导马尔可夫图。此外,观察到交互过程中M2N2存在分割尺寸波动问题,影响整体mIoU。为此,我们将提示建模为序列过程,提出一种新自适应评分函数,综合考虑前次分割结果与当前点击点,防止不合理尺寸变化。以Stable Diffusion 2和Depth Anything V2为骨干网络,实验表明,所提M2N2V2在除医学领域外所有数据集上显著优于M2N2,在点击次数(NoC)和mIoU上均有提升。有趣的是,该无监督方法在更具挑战性的DAVIS和HQSeg44K数据集上,于NoC指标上达到与SAM和SimpleClick等有监督方法相当的性能,缩小了有监督与无监督方法之间的差距。

原文摘要 · Abstract (English)

We present Markov Map Nearest Neighbor V2 (M2N2V2), a novel and simple, yet effective approach which leverages depth guidance and attention maps for unsupervised and training-free point-prompt-based interactive segmentation. Following recent trends in supervised multimodal approaches, we carefully integrate depth as an additional modality to create novel depth-guided Markov-maps. Furthermore, we observe occasional segment size fluctuations in M2N2 during the interactive process, which can decrease the overall mIoU's. To mitigate this problem, we model the prompting as a sequential process and propose a novel adaptive score function which considers the previous segmentation and the current prompt point in order to prevent unreasonable segment size changes. Using Stable Diffusion 2 and Depth Anything V2 as backbones, we empirically show that our proposed M2N2V2 significantly improves the Number of Clicks (NoC) and mIoU compared to M2N2 in all datasets except those from the medical domain. Interestingly, our unsupervised approach achieves competitive results compared to supervised methods like SAM and SimpleClick in the more challenging DAVIS and HQSeg44K datasets in the NoC metric, reducing the gap between supervised and unsupervised methods.

交互分割无监督学习深度图多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。