将传统推荐信号转为软令牌,高效压缩输入提升模型性能
Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models

- 将多种推荐信号转化为可直接输入的软令牌
- 在工业级推荐场景中显著降低提示长度和计算开销
- 适合需要高效融合多源信号的推荐系统研发
大规模推荐模型(LRMs)在工业级推荐任务中展现出良好潜力。然而,如何高效、有效地将传统信号整合进基于Transformer的架构仍是重大挑战。传统方法通过“文本化”或生成离散项表示,常导致提示过长、内存占用大和计算开销高。为此,我们提出“Token Factory”框架,将传统信号转换为可被LRMs直接处理的“软令牌”,实现异构特征的高效集成与压缩,避免提示长度爆炸,同时提升模型性能。我们详述了该框架结构,并在生产规模推荐环境中验证其有效性。
原文摘要 · Abstract (English)
Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically integrating traditional signals into these transformer-based architectures effectively and efficiently remains a major challenge. Conventional approaches that "textualize" these signals directly or create discrete item representations often lead to excessively long prompts, substantial memory footprints, and high computational overhead. To overcome these limitations, we propose "Token Factory", a framework designed to transform traditional signals into "soft tokens" that can be directly processed by LRMs. This approach enables efficient integration and compression of heterogeneous input features, preventing prompt length explosion while enhancing model performance. We detail the architecture of Token Factory and present experimental results validating its effectiveness in a production-scale recommendation environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。