arXiv:2509.24467cs.LGstat.ML2025-09中稿 · ICML

让自监督学习模型可解释,揭示其隐藏偏见。

Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation

  • 用代表性样本点解析模型隐空间,实现无监督可解释性
  • 在百万级数据集上高效计算,支持图像与表格数据
  • 发现收入预测中存在基于人口统计的算法偏见

自监督学习(SSL)从海量无标签数据中学习表征,但模型通常为黑箱,需领域特定解释。我们提出KREPES框架,可解析包括SimCLR、BYOL和VICReg在内的多种SSL目标所学表征。通过将神经网络的实证神经正切核近似与核函数的表示定理结合,我们以“代表性地标”——即关键无标签样本的表征——直接表达学习到的隐空间。引入新指标:样本特异性影响得分、概念条件影响得分与特征对齐差距,量化表征透明度。KREPES无需监督即可审计隐空间,例如在Adult-1M数据集中揭示了模型利用人口统计代理变量进行收入预测的算法偏见。为应对百万级样本基准(ImageNet-1K、Adult-1M),KREPES还提出了基于Nyström近似的分析推断框架,确保可扩展性。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) learns representations from massive unlabeled data, yet the resulting models typically operate as black boxes, necessitating domain-specific explanations. We introduce KREPES, a unified framework to analytically interpret the learned representations of SSL objectives, including SimCLR, BYOL, and VICReg. By bridging empirical neural tangent kernel approximations of neural networks with the Representer Theorem for kernels, we express the learned latent space directly via "Representer Landmarks", which are the representations of influential unlabeled training examples. We introduce novel metrics, "Sample-Specific Influence Score", "Concept-Conditioned Influence Score" and "Feature Alignment Gap", to quantify the transparency of the learned representations. KREPES enables direct audit of the latent space without supervision, for example, revealing an algorithmic bias in the Adult-1M dataset where SSL uses demographic proxies for income. Finally, to ensure scalability to benchmarks with 1M+ samples (ImageNet-1K, Adult-1M), KREPES introduces a novel Nyström approximation-based analytical inference framework for SSL objectives.

自监督学习可解释性偏见检测大规模分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。