arXiv:2501.03469cs.CV2025-01被引 2

通过软变量离散化提升自监督图像表征的可解释性与性能

Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation Learning

  • 用软离散化方法在隐空间估计变量分布,指导信息最大化学习
  • 在多个下游任务中超越对比学习方法,精度与效率双优
  • 结果具备变量级可解释性,适合需要透明模型的场景

自监督学习(SSL)已成为图像处理、编码与理解的关键技术,尤其在无标注大规模数据下构建视觉基础模型方面发挥重要作用。本文提出一种新型自监督方法——信息最大化软变量离散化(IMSVD),通过在隐空间对每个变量进行软离散化,实现对训练批次中变量概率分布的估计,并使学习过程直接由信息度量引导。基于多视角假设,设计了信息论目标函数,以学习变换不变、非冗余、无轨迹的表征特征。进一步推导出用于自监督图像表征学习的联合交叉熵损失,理论上优于现有方法,能有效降低特征冗余。值得注意的是,该非对比性方法在统计上表现如对比学习。大量实验表明,IMSVD在多种下游任务中兼具高准确率与高效率。得益于变量离散化,其优化后的嵌入特征在变量层面具有独特可解释性。该方法可拓展至其他学习范式。代码已公开于 https://github.com/niuchuangnn/IMSVD。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has emerged as a crucial technique in image processing, encoding, and understanding, especially for developing today's vision foundation models that utilize large-scale datasets without annotations to enhance various downstream tasks. This study introduces a novel SSL approach, Information-Maximized Soft Variable Discretization (IMSVD), for image representation learning. Specifically, IMSVD softly discretizes each variable in the latent space, enabling the estimation of their probability distributions over training batches and allowing the learning process to be directly guided by information measures. Motivated by the MultiView assumption, we propose an information-theoretic objective function to learn transform-invariant, non-travail, and redundancy-minimized representation features. We then derive a joint-cross entropy loss function for self-supervised image representation learning, which theoretically enjoys superiority over the existing methods in reducing feature redundancy. Notably, our non-contrastive IMSVD method statistically performs contrastive learning. Extensive experimental results demonstrate the effectiveness of IMSVD on various downstream tasks in terms of both accuracy and efficiency. Thanks to our variable discretization, the embedding features optimized by IMSVD offer unique explainability at the variable level. IMSVD has the potential to be adapted to other learning paradigms. Our code is publicly available at https://github.com/niuchuangnn/IMSVD.

自监督学习表征学习可解释性信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。