arXiv:2601.05845cs.LGstat.ME2026-01

用改进的链接函数让计数数据分解更灵活,提升可解释性

A New Family of Poisson Non-negative Matrix Factorization Methods Using the Shifted Log Link

  • 引入带调节参数的移位对数链接,让成分从相加变为相乘组合
  • 新方法在真实数据上显著改善分解结果的可读性
  • 适合需要灵活建模计数数据结构的研究者

泊松非负矩阵分解(Poisson NMF)是一种广泛用于发现计数数据可解释‘部件’分解的方法。尽管已有多种泊松NMF变体,但现有方法均假设分解中的‘部件’以加法方式组合,这一假设在某些场景下合理,但在其他场景中不适用。本文提出使用移位对数链接函数的泊松NMF,通过一个可调参数,使模型在加法组合(即标准泊松NMF)与乘法组合之间平滑过渡。我们提供基于最大似然的拟合算法,并设计一种近似方法,大幅降低大规模稀疏数据的计算开销(计算复杂度与数据矩阵中非零元素数量成正比)。我们在多个真实数据集上验证了该方法,结果表明链接函数的选择会显著影响分解效果,在某些情况下,采用移位对数链接能显著提升模型的可解释性。

原文摘要 · Abstract (English)

Poisson non-negative matrix factorization (NMF) is a widely used method to find interpretable "parts-based" decompositions of count data. While many variants of Poisson NMF exist, existing methods assume that the "parts" in the decomposition combine additively. This assumption may be natural in some settings, but not in others. Here we introduce Poisson NMF with the shifted-log link function to relax this assumption. The shifted-log link function has a single tuning parameter, and as this parameter varies the model changes from assuming that parts combine additively (i.e., standard Poisson NMF) to assuming that parts combine more multiplicatively. We provide an algorithm to fit this model by maximum likelihood, and also an approximation that substantially reduces computation time for large, sparse datasets (computations scale with the number of non-zero entries in the data matrix). We illustrate these new methods on a variety of real datasets. Our examples show how the choice of link function in Poisson NMF can substantively impact the results, and how in some settings the use of a shifted-log link function may improve interpretability compared with the standard, additive link.

矩阵分解泊松模型可解释性统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。