提出用户行为压缩定律,指导百亿规模下的分词配置优化。
Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

- 发现分词容量与数据规模的对数近似线性关系。
- 在支付宝百亿级数据上验证分词可缓解性能瓶颈。
- 适配大规模用户表征学习,尤其适合工业级推荐系统。
真实工业场景中用户表征学习常通过增加用户数量、行为序列长度和模型规模来扩展。然而现有方法面临两大挑战:(i) 百亿规模原始数据扩展时出现性能增益递减瓶颈,可通过分词缓解;(ii) 缺乏分词配置随数据规模变化的定量分析。本文基于百亿规模支付宝数据集开展初步研究,揭示原始数据扩展瓶颈及分词带来的持续收益。通过理论分析与系统实验,总结出最小必要分词容量与输入数据规模(以token计)的对数间近似线性关系,且斜率随分词方法与数据源变化,反映表示空间冗余与源内独特性的差异。基于该定律,提出ALGN自适应变长分词方法,优化容量分配。跨多种数据源、分词方法与下游任务的实验表明该定律具普适性与可靠性,为大规模用户表征学习中的分词配置提供实践指导,且ALGN优于现有基线。
原文摘要 · Abstract (English)
User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance exhibit diminishing performance gains with larger-scale raw text user behavioral input, which can be mitigated by tokenization. (ii) Lack of quantitative analysis of how tokenization configurations should scale with data size. In this report, we propose User Behavioral Densing Law for characterizing the quantitative relationship between data scale and the minimum sufficient tokenization capacity. Firstly, we conduct a pilot study on raw & tokenized scaling comparison on billion-scale Alipay dataset, revealing the raw data scaling bottleneck and the sustained gains enabled by tokenization. To derive the scaling pattern governing the minimum sufficient tokenization configuration at different data scales, theoretical analysis and systematic experiments are employed to summarize the quantitative scaling pattern. We find an approximately linear relationship between the logarithms of minimum sufficient tokenization capacity and input data size measured by tokens, and the scaling slope varies systematically with the tokenization method and data source, reflecting differences in representation-space redundancy and intra-source uniqueness. Guided by the proposed law, we further develop ALGN, an adaptive variable-length tokenization method that improves capacity allocation. Extensive experiments across diverse data sources, tokenization methods, and downstream tasks demonstrate the generalizability and reliability of the User Behavioral Densing Law, providing practical guidance for tokenization configuration selection in large-scale user representation learning. Moreover, ALGN outperforms existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。