提出可解释的多子空间分析方法,提升高维数据建模的可理解性。
Disentangling Interpretable Factors with Supervised Independent Subspace Principal Component Analysis
- 基于HSIC约束的监督子空间分解,实现多空间解耦
- 在乳腺癌诊断、衰老甲基化等任务中分离出关键生物特征
- 适合需要可解释性的生物医学数据分析场景
机器学习的成功依赖于对高维数据的有效表示,但确保表示结果符合人类可理解概念仍具挑战,常需引入先验知识并分解数据至多个子空间。传统线性方法难以建模多个空间,而更强大的深度学习方法则缺乏可解释性。本文提出监督独立子空间主成分分析(sisPCA),一种扩展自PCA的多子空间学习方法。通过希尔伯特-施密特独立性准则(HSIC)引入监督信号,并同时保证子空间解耦。我们证明了sisPCA与自编码器及正则化线性回归的联系,并在多个应用中展示了其识别和分离隐藏数据结构的能力:包括从图像特征中诊断乳腺癌、学习与衰老相关的DNA甲基化变化,以及疟疾感染的单细胞分析。结果显示与疟疾定植相关的关键功能通路,凸显了可解释表示在高维数据分析中的重要性。
原文摘要 · Abstract (English)
The success of machine learning models relies heavily on effectively representing high-dimensional data. However, ensuring data representations capture human-understandable concepts remains difficult, often requiring the incorporation of prior knowledge and decomposition of data into multiple subspaces. Traditional linear methods fall short in modeling more than one space, while more expressive deep learning approaches lack interpretability. Here, we introduce Supervised Independent Subspace Principal Component Analysis ($\texttt{sisPCA}$), a PCA extension designed for multi-subspace learning. Leveraging the Hilbert-Schmidt Independence Criterion (HSIC), $\texttt{sisPCA}$ incorporates supervision and simultaneously ensures subspace disentanglement. We demonstrate $\texttt{sisPCA}$'s connections with autoencoders and regularized linear regression and showcase its ability to identify and separate hidden data structures through extensive applications, including breast cancer diagnosis from image features, learning aging-associated DNA methylation changes, and single-cell analysis of malaria infection. Our results reveal distinct functional pathways associated with malaria colonization, underscoring the essentiality of explainable representation in high-dimensional data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。