通过精选数据提升生成式推荐系统持续学习效率
Efficient Dataset Selection for Continual Adaptation of Generative Recommenders

- 用梯度特征与分布匹配挑选关键用户交互数据
- 小样本训练下性能下降减少,保持对时序漂移的鲁棒性
- 适合需要频繁更新的线上推荐系统
推荐系统需持续适应用户行为变化,但大规模流式环境中数据量巨大,频繁全量重训不切实际。本文研究如何通过针对性数据选择缓解因时间分布漂移导致的性能下降,同时保持可扩展性。评估了多种表示方式和采样策略,用于构建少量但信息丰富的用户交互数据子集。结果表明,基于梯度的表示结合分布匹配,能有效提升下游模型性能,在实现训练效率提升的同时,维持对分布漂移的鲁棒性。这些发现表明,数据精炼是生产级推荐系统中实现可扩展监控与自适应更新的有效机制。
原文摘要 · Abstract (English)
Recommendation systems must continuously adapt to evolving user behavior, yet the volume of data generated in large-scale streaming environments makes frequent full retraining impractical. This work investigates how targeted data selection can mitigate performance degradation caused by temporal distributional drift while maintaining scalability. We evaluate a range of representation choices and sampling strategies for curating small but informative subsets of user interaction data. Our results demonstrate that gradient-based representations, coupled with distribution-matching, improve downstream model performance, achieving training efficiency gains while preserving robustness to drift. These findings highlight data curation as a practical mechanism for scalable monitoring and adaptive model updates in production-scale recommendation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。