arXiv:2509.15709cs.IR2025-09被引 1

发现推荐系统嵌入维度缩放中的双峰与对数现象,揭示性能变化的深层原因。

Understanding Embedding Scaling in Collaborative Filtering

  • 在10个数据集上测试4种经典模型,系统研究嵌入维度影响。
  • 观察到性能随维度增加出现先升后降再回升的双峰曲线。
  • 理论分析噪声鲁棒性,解释双峰现象并验证于实际数据。

将推荐模型扩展为大规模模型已成为广泛讨论的话题。近期研究关注嵌入维度以外的组件,因普遍认为增大嵌入维度会导致性能下降。尽管已有初步观察,但其不可扩展性的根本原因仍不明确,且性能退化是否普遍存在于不同模型和数据集尚未被探索。本文在10个具有不同稀疏度和规模的数据集上,使用4种代表性经典架构进行大规模实验。令人惊讶地发现了两个新现象:双峰现象(随着嵌入维度增加,性能先提升、后下降、再上升,最终下降)和对数现象(表现呈完美对数曲线)。贡献有三:首先,首次发现协同过滤模型缩放中的两种新现象;其次,揭示了双峰现象的内在机制;最后,从理论上分析了协同过滤模型的噪声鲁棒性,结果与实证观察一致。

原文摘要 · Abstract (English)

Scaling recommendation models into large recommendation models has become one of the most widely discussed topics. Recent efforts focus on components beyond the scaling embedding dimension, as it is believed that scaling embedding may lead to performance degradation. Although there have been some initial observations on embedding, the root cause of their non-scalability remains unclear. Moreover, whether performance degradation occurs across different types of models and datasets is still an unexplored area. Regarding the effect of embedding dimensions on performance, we conduct large-scale experiments across 10 datasets with varying sparsity levels and scales, using 4 representative classical architectures. We surprisingly observe two novel phenomena: double-peak and logarithmic. For the former, as the embedding dimension increases, performance first improves, then declines, rises again, and eventually drops. For the latter, it exhibits a perfect logarithmic curve. Our contributions are threefold. First, we discover two novel phenomena when scaling collaborative filtering models. Second, we gain an understanding of the underlying causes of the double-peak phenomenon. Lastly, we theoretically analyze the noise robustness of collaborative filtering models, with results matching empirical observations.

推荐系统嵌入维度双峰现象协同过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。