arXiv:2501.15816cs.IRcs.AI2025-01中稿 · DASFAA2025被引 1

解决推荐系统长尾数据下特征学习不全面的问题。

AdaF^2M^2: Comprehensive Learning and Responsive Leveraging Features in Recommendation System

  • 通过自适应特征掩码机制,多轮训练增强特征学习
  • 线上测试提升用户活跃天数1.37%、使用时长1.89%
  • 适用于多种推荐场景,对主流模型兼容性强

特征建模(特征表示学习与利用)在工业级推荐系统中至关重要。然而,现实应用中的数据分布通常呈现高度偏斜的长尾模式,导致模型过度依赖用户/物品ID等标识特征,难以充分学习非ID类元特征(如用户/物品属性)。这不仅限制了特征学习的全面性,也削弱了模型的泛化能力,使其更易受数据噪声影响。现有研究多关注特征提取与交互,忽视长尾分布带来的问题。为此,我们提出一种模型无关框架AdaF^2M^2(自适应特征建模与特征掩码),通过特征掩码机制实现多轮训练下的增强学习,结合适配器对不同用户/物品状态动态加权特征。在多个推荐场景的在线A/B测试中,分别取得用户活跃天数+1.37%、应用时长+1.89%的提升。离线实验亦验证其在不同模型上的有效性。该方法已在抖音集团多个应用的召回与排序任务中广泛部署,展现出卓越效果与普适性。

原文摘要 · Abstract (English)

Feature modeling, which involves feature representation learning and leveraging, plays an essential role in industrial recommendation systems. However, the data distribution in real-world applications usually follows a highly skewed long-tail pattern due to the popularity bias, which easily leads to over-reliance on ID-based features, such as user/item IDs and ID sequences of interactions. Such over-reliance makes it hard for models to learn features comprehensively, especially for those non-ID meta features, e.g., user/item characteristics. Further, it limits the feature leveraging ability in models, getting less generalized and more susceptible to data noise. Previous studies on feature modeling focus on feature extraction and interaction, hardly noticing the problems brought about by the long-tail data distribution. To achieve better feature representation learning and leveraging on real-world data, we propose a model-agnostic framework AdaF^2M^2, short for Adaptive Feature Modeling with Feature Mask. The feature-mask mechanism helps comprehensive feature learning via multi-forward training with augmented samples, while the adapter applies adaptive weights on features responsive to different user/item states. By arming base models with AdaF^2M^2, we conduct online A/B tests on multiple recommendation scenarios, obtaining +1.37% and +1.89% cumulative improvements on user active days and app duration respectively. Besides, the extended offline experiments on different models show improvements as well. AdaF$^2$M$^2$ has been widely deployed on both retrieval and ranking tasks in multiple applications of Douyin Group, indicating its superior effectiveness and universality.

推荐系统特征学习长尾分布自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。