梳理五类非负矩阵分解模型的共性与可识别性问题
A review of NMF, PLSA, LBA, EMA, and LCA with a focus on the identifiability issue
- 揭示LBA、LCA、EMA、PLSA与NMF本质相同
- 证明五类模型解的唯一性等价于NMF的唯一性
- 适合关注模型可解释性与理论基础的研究者
在机器学习、社会科学、地理学等领域,非负矩阵分解及其变体被广泛用于将非负矩阵分解为两个或三个矩阵的乘积,受限于非负性或行和为1的约束。尽管这些模型在很大程度上相似甚至等价,但它们以不同名称出现,其内在联系未被充分认知。本文重点分析五种流行模型——潜预算分析(LBA)、潜类别分析(LCA)、端元分析(EMA)、概率潜在语义分析(PLSA)与非负矩阵分解(NMF)之间的相似性。特别聚焦于模型的可识别性问题,证明了LBA、EMA、LCA、PLSA的解唯一当且仅当NMF的解唯一。文章还简要回顾了各类模型的算法,并以社会科学研究中的时间预算数据集为例进行说明,最后讨论了与弧线分析等密切相关模型的关系。
原文摘要 · Abstract (English)
Across fields such as machine learning, social science, geography, considerable attention has been given to models that factorize a nonnegative matrix into the product of two or three matrices, subject to nonnegative or row-sum-to-1 constraints. Although these models are to a large extend similar or even equivalent, they are presented under different names, and their similarity is not well known. This paper highlights similarities among five popular models, latent budget analysis (LBA), latent class analysis (LCA), end-member analysis (EMA), probabilistic latent semantic analysis (PLSA), and nonnegative matrix factorization (NMF). We focus on an essential issue-identifiability-of these models and prove that the solution of LBA, EMA, LCA, PLSA is unique if and only if the solution of NMF is unique. We also provide a brief review for algorithms of these models. We illustrate the models with a time budget dataset from social science, and end the paper with a discussion of closely related models such as archetypal analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。