千亿级用户场景下统一建模兴趣,提升推荐精准度。
Cross-Scenario Unified Modeling of User Interests at Billion Scale
- 用双塔LLM框架融合多场景行为信号,实现统一兴趣表示。
- 在线测试覆盖超亿用户,推荐与广告效果显著提升。
- 适合大规模内容平台构建跨场景个性化推荐系统。
内容平台上的用户兴趣具有天然多样性,表现为搜索、信息流浏览和内容发现等异构场景中的复杂行为模式。传统推荐系统通常只关注单一场景的业务指标优化,忽视跨场景行为信号,且难以在千亿规模部署中整合大模型技术,限制了对全链路用户兴趣的捕捉能力。本文提出RED-Rec——一种面向工业级内容推荐的增强型分层推荐引擎,通过聚合与合成多场景行为动作,实现跨场景用户兴趣表征的统一,形成全面的物品与用户建模。其核心采用双塔式大模型架构,在保证部署效率的同时生成细粒度、多维度的表示,并设计场景感知的密集混合与查询策略,有效融合多样化行为信号,捕捉跨场景用户意图模式并支持服务阶段的细粒度上下文表达。我们在红笔记(RedNote)平台对数亿用户进行在线A/B测试,验证了其在内容推荐与广告投放任务中的显著性能提升。此外,我们发布了百万级序列推荐数据集RED-MMU,用于离线训练与评估。本工作推动了统一用户建模的发展,为大规模UGC平台实现深度个性化与更深层次用户互动提供了可能。
原文摘要 · Abstract (English)
User interests on content platforms are inherently diverse, manifesting through complex behavioral patterns across heterogeneous scenarios such as search, feed browsing, and content discovery. Traditional recommendation systems typically prioritize business metric optimization within isolated specific scenarios, neglecting cross-scenario behavioral signals and struggling to integrate advanced techniques like LLMs at billion-scale deployments, which finally limits their ability to capture holistic user interests across platform touchpoints. We propose RED-Rec, an LLM-enhanced hierarchical Recommender Engine for Diversified scenarios, tailored for industry-level content recommendation systems. RED-Rec unifies user interest representations across multiple behavioral contexts by aggregating and synthesizing actions from varied scenarios, resulting in comprehensive item and user modeling. At its core, a two-tower LLM-powered framework enables nuanced, multifaceted representations with deployment efficiency, and a scenario-aware dense mixing and querying policy effectively fuses diverse behavioral signals to capture cross-scenario user intent patterns and express fine-grained, context-specific intents during serving. We validate RED-Rec through online A/B testing on hundreds of millions of users in RedNote through online A/B testing, showing substantial performance gains in both content recommendation and advertisement targeting tasks. We further introduce a million-scale sequential recommendation dataset, RED-MMU, for comprehensive offline training and evaluation. Our work advances unified user modeling, unlocking deeper personalization and fostering more meaningful user engagement in large-scale UGC platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。