arXiv:2509.21149cs.LGcs.AI2025-09

让无监督嵌入可解释,发现输入特征的局部关联模式

LAVA: Explainability for Unsupervised Latent Embeddings

  • 通过输入数据特征共变关系揭示嵌入局部结构
  • 能稳定生成多粒度解释,捕捉全局重复的子模式
  • 适合图像部件、细胞疾病信号等真实场景解释

无监督黑箱模型推动科学发现,但其输出为多维嵌入,难以解释。现有解释方法或仅针对单样本,或仅提供数据整体摘要,过于琐碎或片面,且无法在无映射函数下解释嵌入。为此,我们提出LAVA,一种后置、模型无关的方法,通过原始输入数据中特征的共变关系解释嵌入的局部组织结构。LAVA解释由模块组成,捕捉在嵌入中反复出现的输入特征相关性局部子模式。该方法能在指定粒度下生成稳定解释,揭示图像视觉部件或细胞过程中的疾病信号等领域相关模式,现有方法难以发现。

原文摘要 · Abstract (English)

Unsupervised black-box models are drivers of scientific discovery, yet are difficult to interpret, as their output is often a multidimensional embedding rather than a well-defined target. While explainability for supervised learning uncovers how input features contribute to predictions, its unsupervised counterpart should relate input features to the structure of the learned embeddings. However, adaptations of supervised model explainability for unsupervised learning provide either single-sample or dataset-summary explanations, remaining too fine-grained or reductive to be meaningful, and cannot explain embeddings without mapping functions. To bridge this gap, we propose LAVA, a post-hoc model-agnostic method to explain local embedding organization through feature covariation in the original input data. LAVA explanations comprise modules, capturing local subpatterns of input feature correlation that reoccur globally across the embeddings. LAVA delivers stable explanations at a desired level of granularity, revealing domain-relevant patterns such as visual parts of images or disease signals in cellular processes, otherwise missed by existing methods.

可解释性无监督学习嵌入解释特征共变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。