通过最大化基向量体积提升NMF可识别性
Identification of NMF by choosing maximum-volume basis vectors
- 新框架强制基向量尽可能分离,避免混合
- 理论证明可唯一识别真实基向量
- 适合高混合度数据,结果更易解释
非负矩阵分解(NMF)中,最小体积约束NMF通过使基向量尽可能紧凑来识别解,通常导致系数矩阵稀疏(每行含零元素)。然而,对于高度混合的数据,这种稀疏性不成立,且估计的基向量可能为真实基向量的混合,难以解释。为此,本文提出最大体积约束NMF框架,使基向量尽可能分离。我们建立了该方法的可识别性定理,并提供了相应的算法。实验结果验证了该方法的有效性。
原文摘要 · Abstract (English)
In nonnegative matrix factorization (NMF), minimum-volume-constrained NMF is a widely used framework for identifying the solution of NMF by making basis vectors as similar as possible. This typically induces sparsity in the coefficient matrix, with each row containing zero entries. Consequently, minimum-volume-constrained NMF may fail for highly mixed data, where such sparsity does not hold. Moreover, the estimated basis vectors in minimum-volume-constrained NMF may be difficult to interpret as they may be mixtures of the ground truth basis vectors. To address these limitations, in this paper we propose a new NMF framework, called maximum-volume-constrained NMF, which makes the basis vectors as distinct as possible. We further establish an identifiability theorem for maximum-volume-constrained NMF and provide an algorithm to estimate it. Experimental results demonstrate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。