针对高维稀疏共现数据,提出共享参数零膨胀伽马模型提升预测精度。
Matrix factorization and prediction for high dimensional co-occurrence count data via shared parameter alternating zero inflated Gamma model
- 用零膨胀伽马分布建模共现频次,通过向量余弦相似度刻画项目相关性。
- 引入学习率调整的迭代算法,在模拟中稳定收敛并准确估计参数。
- 适用于电商推荐、词对共现等高维稀疏数据场景,尤其适合零值过多的情况。
高维稀疏矩阵数据广泛存在于各类应用中。典型例子是加权词-词共现计数数据,统计词语对在同一上下文窗口内出现的加权频率,这类数据通常具有高度偏斜的非负值且包含大量零值。另一个例子是电商中的物品-物品或用户-物品共现数据。目标是利用此类数据预测项目或用户间的相关性。本文假设项目或用户可由未知稠密向量表示,将共现计数视为来自零膨胀伽马随机变量,并使用未知向量间的余弦相似度来总结项目-项目相关性。采用共享参数交替零膨胀伽马回归模型(SA-ZIG)估计未知值。考虑了标准链接和对数链接两种形式。提出了两种参数更新方案,并给出了参数估计算法。进行了收敛性分析。数值研究显示,使用Fisher scoring但不调整学习率的SA-ZIG可能无法找到最大似然估计;而采用学习率调整的SA-ZIG在模拟中表现良好。
原文摘要 · Abstract (English)
High-dimensional sparse matrix data frequently arise in various applications. A notable example is the weighted word-word co-occurrence count data, which summarizes the weighted frequency of word pairs appearing within the same context window. This type of data typically contains highly skewed non-negative values with an abundance of zeros. Another example is the co-occurrence of item-item or user-item pairs in e-commerce, which also generates high-dimensional data. The objective is to utilize this data to predict the relevance between items or users. In this paper, we assume that items or users can be represented by unknown dense vectors. The model treats the co-occurrence counts as arising from zero-inflated Gamma random variables and employs cosine similarity between the unknown vectors to summarize item-item relevance. The unknown values are estimated using the shared parameter alternating zero-inflated Gamma regression models (SA-ZIG). Both canonical link and log link models are considered. Two parameter updating schemes are proposed, along with an algorithm to estimate the unknown parameters. Convergence analysis is presented analytically. Numerical studies demonstrate that the SA-ZIG using Fisher scoring without learning rate adjustment may fail to fi nd the maximum likelihood estimate. However, the SA-ZIG with learning rate adjustment performs satisfactorily in our simulation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。