让推荐系统高效又精准地理解用户长期行为
SITA: Semantic Interest Tokens for Target-Aware Compression in Long-Sequence Recommendation

- 用语义标识压缩用户行为,构建结构化兴趣
- 根据目标物品动态聚合兴趣,实现精准匹配
- 兼顾效率与适应性,适合大规模推荐场景
随着互联网平台用户行为序列不断增长,有效建模长序列行为对预测用户兴趣至关重要。现有方法分为两类:一类通过动态检索相关行为实现目标感知建模,但推理时需依赖目标计算;另一类将完整行为序列压缩为紧凑用户表示,虽高效可扩展,却因目标无关编码牺牲了目标自适应能力。核心挑战在于如何在保持压缩表示效率的同时实现目标感知建模。为此,我们提出SITA——一种面向长序列推荐的目标感知压缩框架。SITA通过并行语义量化学习语义标识,将压缩兴趣组织成语义结构;基于目标物品的语义标识,自适应聚合对应结构化兴趣,构建目标特定用户表示。在公开数据集和大规模工业数据集上的实验表明,SITA持续优于代表性基线,同时保持强可扩展性,展现出在真实推荐系统中的巨大潜力。
原文摘要 · Abstract (English)
As user behavior histories continue to grow on modern Internet platforms, effectively modeling long behavior sequences has become crucial for predicting user interests in candidate items. Existing methods have evolved along two directions. One line dynamically retrieves target-relevant behaviors from long histories, enabling target-aware modeling but requiring target-dependent computation during inference. The other line compresses entire behavior sequences into compact user representations, achieving high efficiency and scalability but sacrificing target-specific adaptation due to target-independent encoding. The key challenge is therefore to enable target-aware modeling while preserving the efficiency and scalability of compressed user representations. To address this challenge, we propose \textbf{SITA}, a target-aware compression framework for long-sequence recommendation. SITA enables target-aware compression by organizing compressed interests into semantic structures through semantic identifiers learned via parallel semantic quantization. Conditioned on the semantic identifier of the target item, SITA adaptively aggregates the corresponding structured interests to construct the target-specific user representation. Extensive experiments on public datasets and a large-scale industrial dataset demonstrate that SITA consistently outperforms representative baselines while maintaining strong scalability, highlighting its strong potential for real-world recommender systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。