arXiv:2409.06589cs.CV2024-09

用双曲图网络实现轻量级无监督图像分割,效果超越现有方法。

Seg-HGNN: Unsupervised and Light-Weight Image Segmentation with Hyperbolic Graph Neural Networks

  • 基于双曲空间构建轻量图神经网络,以小尺寸嵌入捕捉图像层次结构。
  • 在VOC-07/VOC-12上定位精度提升2.5%/4%,在CUB-200/ECSSD上分割性能提升0.8%/1.3%。
  • 参数少于7.5k,在GTX1650上每秒处理约2张图,适合资源受限场景。

在欧氏空间中通过线性超空间进行图像分析已有深入研究。然而,为获得更有效的图像表示,我们转向双曲流形。它们能以极低维度有效捕捉图像中的复杂层级关系。为验证双曲嵌入的能力,我们提出一种轻量级双曲图神经网络用于图像分割,以极小嵌入尺寸整合补丁级特征。所提方法Seg-HGNN在未标注数据下表现优于当前最佳无监督方法:在VOC-07和VOC-12上的定位精度分别提升2.5%和4%,在CUB-200和ECSSD上的分割性能分别提升0.8%和1.3%。模型参数少于7.5k,可在如GTX1650等标准显卡上以约2张/秒的速度运行。实证结果充分证明了双曲表示在视觉任务中的有效性与潜力。

原文摘要 · Abstract (English)

Image analysis in the euclidean space through linear hyperspaces is well studied. However, in the quest for more effective image representations, we turn to hyperbolic manifolds. They provide a compelling alternative to capture complex hierarchical relationships in images with remarkably small dimensionality. To demonstrate hyperbolic embeddings' competence, we introduce a light-weight hyperbolic graph neural network for image segmentation, encompassing patch-level features in a very small embedding size. Our solution, Seg-HGNN, surpasses the current best unsupervised method by 2.5\%, 4\% on VOC-07, VOC-12 for localization, and by 0.8\%, 1.3\% on CUB-200, ECSSD for segmentation, respectively. With less than 7.5k trainable parameters, Seg-HGNN delivers effective and fast ($\approx 2$ images/second) results on very standard GPUs like the GTX1650. This empirical evaluation presents compelling evidence of the efficacy and potential of hyperbolic representations for vision tasks.

图像分割双曲几何轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。