用精心设计的合成数据,首次让推荐大模型出现可预测的缩放规律。
Principled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendation
- 构建分层合成数据框架,通过教学式课程提升数据质量。
- 在召回率@100上,模型性能比真实数据训练高出130%。
- 适合关注推荐系统大模型可扩展性的研究者和工程师。
大语言模型(LLM)为推荐系统带来新可能,但其发展受限于缺乏可预测的缩放规律,而这一问题可能源于以往持续预训练中原始用户行为数据的噪声、偏差与不完整性。本文提出一种新型分层合成数据生成框架,通过构建结构化、有指导性的教学课程,规避上述问题。实验表明,基于该合成数据训练的标准序列模型,在下游排序任务中显著优于真实数据训练的模型,如SasRec在recall@100上提升130%。更重要的是,本文首次在推荐领域实证了持续预训练大模型的稳健幂律缩放关系:多种合成数据模态下,困惑度均呈现一致且可预测的下降趋势。这确立了推荐领域可靠缩放大模型能力的基础方法,推动研究重心从修复数据缺陷转向利用高质量结构化信息。
原文摘要 · Abstract (English)
Large Language Models (LLMs) represent a promising frontier for recommender systems, yet their development has been impeded by the absence of predictable scaling laws, which are crucial for guiding research and optimizing resource allocation. We hypothesize that this may be attributed to the inherent noise, bias, and incompleteness of raw user interaction data in prior continual pre-training (CPT) efforts. This paper introduces a novel, layered framework for generating high-quality synthetic data that circumvents such issues by creating a curated, pedagogical curriculum for the LLM. We provide powerful, direct evidence for the utility of our curriculum by showing that standard sequential models trained on our principled synthetic data significantly outperform ($+130\%$ on recall@100 for SasRec) models trained on real data in downstream ranking tasks, demonstrating its superiority for learning generalizable user preference patterns. Building on this, we empirically demonstrate, for the first time, robust power-law scaling for an LLM that is continually pre-trained on our high-quality, recommendation-specific data. Our experiments reveal consistent and predictable perplexity reduction across multiple synthetic data modalities. These findings establish a foundational methodology for reliable scaling LLM capabilities in the recommendation domain, thereby shifting the research focus from mitigating data deficiencies to leveraging high-quality, structured information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。