arXiv:2501.00623cs.LGstat.ML2025-01被引 1

提出新模型用共享参数建模高维共现数据,适合推荐系统和文本分析。

Global dense vector representations for words or items using shared parameter alternating Tweedie model

  • 用加权窗口提取共现数值,构建共享参数交替Tweedie模型
  • 结合学习率调整的Fisher评分法在模拟与应用中表现最优
  • 适用于大规模共现数据,尤其适合无观测协变量场景

本文针对在线购物平台的用户-项目或项目-项目共现数据、文本序列中的词-词共现对等实际场景中的共现计数数据,提出一种分析方法。传统回归模型不适用,因缺乏协变量观测,且共现矩阵维度极高,难以放入内存。我们通过定义基于连续尺度加权计数的共现窗口,提取数值数据,并允许零观测具有正概率质量。提出共享参数交替Tweedie(SA-Tweedie)模型及参数估计算法,引入学习率调整配合内层Fisher评分法以确保优化方向稳定。也考虑了使用Adam更新的梯度下降作为替代方法。模拟研究与实际应用表明,采用学习率调整的Fisher评分法优于其他两种方法。同时研究了伪似然方法与交替参数更新,但数值研究表明该方法在无观测协变量的共享参数交替回归模型中不适用。

原文摘要 · Abstract (English)

In this article, we present a model for analyzing the cooccurrence count data derived from practical fields such as user-item or item-item data from online shopping platform, cooccurring word-word pairs in sequences of texts. Such data contain important information for developing recommender systems or studying relevance of items or words from non-numerical sources. Different from traditional regression models, there are no observations for covariates. Additionally, the cooccurrence matrix is typically of so high dimension that it does not fit into a computer's memory for modeling. We extract numerical data by defining windows of cooccurrence using weighted count on the continuous scale. Positive probability mass is allowed for zero observations. We present Shared parameter Alternating Tweedie (SA-Tweedie) model and an algorithm to estimate the parameters. We introduce a learning rate adjustment used along with the Fisher scoring method in the inner loop to help the algorithm stay on track of optimizing direction. Gradient descent with Adam update was also considered as an alternative method for the estimation. Simulation studies and an application showed that our algorithm with Fisher scoring and learning rate adjustment outperforms the other two methods. Pseudo-likelihood approach with alternating parameter update was also studied. Numerical studies showed that the pseudo-likelihood approach is not suitable in our shared parameter alternating regression models with unobserved covariates.

共现建模高维数据推荐系统Tweedie模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。