arXiv:2604.25839cs.IR2026-04

用用户转化后内容提升留存预测,避免数据泄露。

Break the Inaccessible Boundary: Distilling Post-Conversion Content for User Retention Modeling

论文配图:Break the Inaccessible Boundary: Distilling Post-Conversion Content for User Retention Modeling
图 1 · 摘自论文原文
  • 分两阶段蒸馏,让模型在不看未来内容时也能学好留存信号。
  • 线上实验显示留存率显著提升,真实业务场景效果稳定。
  • 适合做广告推荐、用户增长的工程师和研究者参考。

用户留存是衡量平台长期活跃的关键指标。在实时竞价(RTB)广告系统中,为实现用户召回,留存模型需在用户转化前的竞价时刻预测其未来回访概率。尽管转化后的“引导内容”(Onboarding Content)对留存预测具有高度信息价值,但直接用于训练会导致严重的特征泄露,造成训练与推理阶段的差异。为此,我们提出OCARM框架——一种面向引导内容增强的留存建模两阶段蒸馏对齐方法,使模型在推理时仅依赖可观测特征即可隐式捕捉未来内容信号。第一阶段,通过显式暴露引导内容训练一个分层编码器生成教师表示;第二阶段,用户编码器通过蒸馏与冻结的教师对齐,从而在无泄露的前提下逼近不可见的引导信号。大量离线实验与在线A/B测试表明,该框架在真实增长场景中实现了持续性能提升。

原文摘要 · Abstract (English)

User retention is a key metric to measure long-term engagement in modern platforms. In real-time bidding (RTB) advertising system for user re-engagement, the retention model is required to predict future revisit probability at bidding time, before the user converts and consumes any content. Although post-conversion content, termed Onboarding Content, provides highly informative signals for retention prediction, directly using it in training causes severe feature leakage and creates a gap between training and serving. To address this issue, we propose OCARM, a two-stage distillation-aligned framework for Onboarding Content Augmented Retention Modeling, enabling the model to implicitly capture future content using only observable features during inference. In the first stage, we deliberately expose onboarding content to train a hierarchical encoder that produces teacher representations. In the second stage, a user encoder is aligned with the frozen teacher through distillation, allowing the model to approximate the inaccessible onboarding signals without leakage. Extensive offline experiments and online A/B tests demonstrate that our framework achieves consistent improvements in a real-world growth scenario.

用户留存模型蒸馏广告系统特征泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。