提出基于球面密度建模的自监督学习方法,提升图像分割性能。
Self-Supervised Representation Learning via Hyperspherical Density Shaping
- 在球面空间中通过非参数冯·米塞斯-费舍尔估计器优化互信息
- 模型更关注图像前景特征,在PASCAL VOC上分割表现优异
- 揭示隐空间几何与学习动态,为理论化自监督学习提供新思路
现代自监督表示学习方法常依赖缺乏理论基础的启发式设计。本文提出一种理论根基扎实的方法HyDeS,基于多视角互信息最大化,在超球面空间中利用香农微分熵与非参数冯·米塞斯-费舍尔密度估计器进行建模。实验表明,HyDeS能引导模型聚焦图像前景特征,在分割任务如PASCAL VOC上表现良好,但在细粒度分类任务上表现较弱。本文还对诱导出的隐空间几何结构与学习动态进行了详细分析,可为设计其他理论驱动的自监督学习方法提供参考。
原文摘要 · Abstract (English)
Modern self-supervised representation learning methods often relies on empirical heuristics that are not theoretically grounded. In this study we propose HyDeS, a theoretically grounded method based on multi-view mutual information maximization within an hyperspherical space using Shannon differential entropy with a non-parametric von Mises-Fisher density estimator. We show that HyDeS bias the trained model towards focusing on foreground features of the images and perform well on segmentation tasks such as VOC PASCAL, while it lags in fine-grained classification. We provide a detailed analysis of the induced latent space geometry and learning dynamics, that can be used for designing other theoretically grounded self-supervised learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。