提出一种可解释的矩阵优化方法,提升高维数据的可读性。
Non-Negative Stiefel Approximating Flow: Orthogonalish Matrix Optimization for Interpretable Embeddings
- 通过连续平衡重建精度与列间去相关性,实现结构化稀疏
- 在白血病和阿尔茨海默病数据上保持或超越现有方法性能
- 适合需要可解释性的生物医学数据分析场景
可解释表示学习是现代机器学习的核心挑战,尤其在神经影像、基因组学和文本分析等高维场景中。现有方法难以兼顾可解释性与模型灵活性,限制了从复杂数据中提取有意义洞见的能力。本文提出非负Stiefel近似流(NSA-Flow),一个通用矩阵估计框架,融合稀疏矩阵分解、正交化与约束流形学习思想。NSA-Flow通过单个可调权重控制重建保真度与列间去相关性的连续平衡,实现结构化稀疏。该方法在Stiefel流形附近以平滑流形式运行,结合近端更新保证非负性与自适应梯度控制,生成兼具稀疏性、稳定性与可解释性的表示。相比传统正则化,NSA-Flow提供直观的全局结构级稀疏操控机制,简化潜在特征。实验证明其目标可平稳优化,并能无缝集成至降维流程,在模拟与真实生物医学数据中提升可解释性与泛化能力。在Golub白血病数据集和阿尔茨海默病研究中,NSA-Flow在几乎无额外方法成本下保持或优于对比方法性能。该方法为可解释机器学习提供了一种可扩展、通用的工具,适用于多领域数据科学。
原文摘要 · Abstract (English)
Interpretable representation learning is a central challenge in modern machine learning, particularly in high-dimensional settings such as neuroimaging, genomics, and text analysis. Current methods often struggle to balance the competing demands of interpretability and model flexibility, limiting their effectiveness in extracting meaningful insights from complex data. We introduce Non-negative Stiefel Approximating Flow (NSA-Flow), a general-purpose matrix estimation framework that unifies ideas from sparse matrix factorization, orthogonalization, and constrained manifold learning. NSA-Flow enforces structured sparsity through a continuous balance between reconstruction fidelity and column-wise decorrelation, parameterized by a single tunable weight. The method operates as a smooth flow near the Stiefel manifold with proximal updates for non-negativity and adaptive gradient control, yielding representations that are simultaneously sparse, stable, and interpretable. Unlike classical regularization schemes, NSA-Flow provides an intuitive geometric mechanism for manipulating sparsity at the level of global structure while simplifying latent features. We demonstrate that the NSA-Flow objective can be optimized smoothly and integrates seamlessly with existing pipelines for dimensionality reduction while improving interpretability and generalization in both simulated and real biomedical data. Empirical validation on the Golub leukemia dataset and in Alzheimer's disease demonstrate that the NSA-Flow constraints can maintain or improve performance over related methods with little additional methodological effort. NSA-Flow offers a scalable, general-purpose tool for interpretable ML, applicable across data science domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。