arXiv:2412.16247cs.LGcs.AI2024-12ICML被引 6

用稀疏字典学习从显微图像模型中挖掘生物概念

Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation models

  • 结合主成分分析去相关预处理与迭代字典学习
  • 成功提取出细胞类型、基因干扰等生物概念
  • 适合关注生物图像可解释性研究的科研人员

稀疏字典学习(DL)已成为从主要在文本领域训练的大语言模型内部提取语义概念的有效方法。本文探讨了该方法能否用于提取较少人类可解释的科学数据中的有意义概念,例如在细胞显微图像上训练的视觉基础模型。我们提出一种新方法,将稀疏字典学习算法与基于控制数据的PCA去相关预处理相结合。通过这一组合,成功提取出细胞类型、基因扰动等生物上合理的概念。此外,该方法揭示了人类可理解干预所引起的细微形态变化,为通过机制可解释性推动生物成像领域的科学发现提供了新方向。

原文摘要 · Abstract (English)

Sparse dictionary learning (DL) has emerged as a powerful approach to extract semantically meaningful concepts from the internals of large language models (LLMs) trained mainly in the text domain. In this work, we explore whether DL can extract meaningful concepts from less human-interpretable scientific data, such as vision foundation models trained on cell microscopy images, where limited prior knowledge exists about which high-level concepts should arise. We propose a novel combination of a sparse DL algorithm, Iterative Codebook Feature Learning (ICFL), with a PCA whitening pre-processing step derived from control data. Using this combined approach, we successfully retrieve biologically meaningful concepts, such as cell types and genetic perturbations. Moreover, we demonstrate how our method reveals subtle morphological changes arising from human-interpretable interventions, offering a promising new direction for scientific discovery via mechanistic interpretability in bioimaging.

字典学习生物成像可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。