揭示推荐系统中用户嵌入因流行度偏见而收敛的数学机制。
Between-User Collapse Under Popularity-Biased Feedback: A Centered-Covariance Theorem and Computable Phase Boundary
- 用中心化协方差分析用户间差异,发现其随训练趋于噪声水平。
- 推导出可计算的强收缩相界,验证在MovieLens-25M上预测准确。
- 实测表明部署时收缩效应小且不影响推荐指标,干预无效。
我们研究了流行度偏见下的BPR训练如何重塑协同过滤嵌入的用户间几何结构。采用均值中心化的用户协方差矩阵 $C = frac{1}{n} U^ op H U$ 来衡量用户之间的可区分性,而非以往使用的未中心化二阶矩。证明在平稳物品假设下,$C$ 收敛至与项目噪声协方差 $Q$ 成比例的稳态,导致用户间差异坍缩至噪声基线。我们推导出一个闭式、可计算的相边界,以训练超参数 $(α, λ_{neg}, γ, d)$ 分离收缩与扩张区域,并在MovieLens-25M上验证了方向性预测。进一步分析发现:在部署级正则化下,预测的收缩虽真实且由策略驱动,但幅度极小,未体现在任何推荐指标中;$α$ 驱动的各向异性坍缩仅出现在损害推荐性能的正则化强度下;基于理论设计的部署时恢复干预未能提升推荐质量。该相边界可通过训练后模型嵌入、项目交互次数及训练超参数直接计算,使从业者无需模拟反馈环即可判断系统是否处于强坍缩区。实验表明,实际部署设置远低于该临界区。
原文摘要 · Abstract (English)
We study how popularity-biased BPR training reshapes the between-user geometry of collaborative-filtering embeddings. We work with the mean-centered user covariance $C=\tfrac1n U^\top H U$, the object that measures how distinguishable users are from one another, as opposed to the uncentered second moment used in prior work. We prove that under popularity-biased feedback with stationary items, $C$ converges to a steady state proportional to the item-noise covariance $Q$. Thus between-user spread collapses toward a noise floor. We derive a closed-form, computable phase boundary in the training hyperparameters $(α,λ_{neg},γ,d)$ separating contraction from expansion, and validate both directional predictions on MovieLens-25M. We then examine the limits of the effect. At deployment-scale regularization the predicted contraction is real and policy-driven but small, and it is not reflected in any recommendation-level metric we measured. The $α$-driven anisotropic-collapse mechanism operates only at regularization strengths that degrade the recommender. A deployment-time restoration intervention derived from the theory does not improve recommendation quality. The boundary is computable from a trained model's embeddings, item interaction counts, and training hyperparameters, so a practitioner can check whether a deployed system sits in the strong-collapse regime without simulating the feedback loop. In our experiments the boundary places deployable settings far from that regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。