发现推荐系统分数中存在一种普遍的热度成分,与模型架构无关。
A Rank-One Popularity Component in Dot-Product Recommender Scores:Population Theory and Prior-Separation Evidence
- 从条件训练分布出发,揭示点积解码器的理论分解机制。
- 实验证明分离热度项可使热度对齐得分能量降低98.6%。
- 适用于理解非Transformer模型中的推荐退化现象。
推荐系统中的表示各向异性常归因于Transformer架构。我们发现更普遍的来源在于条件训练分布本身。对于任意使用点积Softmax解码器的编码器,其总体最优得分可分解为互信息项、物品边际项log p(i)和上下文相关偏置项。中心化后,物品边际项产生共享的秩一得分分量,而随时间变化的边际项则诱导出低秩热度子空间。这一得分层面的结果并不意味着普遍的嵌入坍塌,因其向嵌入的转移依赖因子分解几何结构。在合成数据及公开的Alibaba与Tianchi交互日志上的实验支持该机制。在匹配干预中,将log p(i)从学习到的点积中分离,使测得的热度对齐得分能量下降98.6%。置换检验确认该降幅特异性地对应于实际热度方向。这些结果将一类看似表示退化的现象解释为长尾物品边际导致的解码器级后果,而非仅限于Transformer编码器的特性。
原文摘要 · Abstract (English)
Representation anisotropy in recommender systems is often attributed to Transformer architectures. We identify a more general source in the conditional training distribution. For any encoder using a dot-product softmax decoder, the population-optimal score decomposes into pointwise mutual information, an item-marginal term log p(i), and a context-dependent offset. After centering, the item marginal produces a context-shared rank-one score component, while time-varying marginals induce a low-rank popularity subspace. This score-level result does not imply universal embedding collapse because its transfer to embeddings depends on factorization geometry. Experiments on synthetic data and public Alibaba and Tianchi interaction logs support the proposed mechanism. Separating log p(i) from the learned dot product reduces the measured popularity-aligned score energy by 98.6 percent in a matched intervention. Permutation tests confirm that this reduction is specific to the empirical popularity direction. These results explain a class of apparent representation degeneration as a decoder-level consequence of long-tailed item marginals rather than a property unique to Transformer encoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。