用决策变压器优化通知推送,提升相关性并减少用户疲劳。
Generative Sequential Notification Optimization via Multi-Objective Decision Transformers
- 将通知决策转为基于回报的监督学习,增强稳定性与可扩展性。
- 在领英系统中使会话量提升0.72%,同时降低用户疲劳度。
- 适合需要高可靠性、实时响应的通知系统优化场景。
通知是传递及时、相关资讯的重要渠道。优化通知推送涉及在信息效用与用户疲劳等约束下的复杂序列决策问题。现有离线强化学习方法(如保守Q学习,CQL)虽被应用,但在大规模场景下存在不稳定性、对分布偏移敏感、可复现性差及高维推荐场景下解释性不足等挑战。本文提出基于决策变压器(Decision Transformer, DT)的框架,将策略学习重构为回报条件化的监督学习,提升了鲁棒性、可扩展性和建模灵活性。主要贡献包括:与CQL的真实世界对比实验;适用于非周期任务的多奖励设计;基于分位数回归的回报预期条件化方法;以及采用环形缓冲区的生产级系统,支持近实时推理。大量离线与在线实验表明,该方法在提升通知效用和整体会话活跃度的同时有效控制用户疲劳。相较于多目标CQL代理,在领英系统中使通知决策的会话量提升0.72%。
原文摘要 · Abstract (English)
Notifications are an important communication channel for delivering timely and relevant information. Optimizing their delivery involves addressing complex sequential decision-making challenges under constraints such as message utility and user fatigue. Offline reinforcement learning (RL) methods, such as Conservative Q-Learning (CQL), have been applied to this problem but face practical challenges at scale, including instability, sensitivity to distribution shifts, limited reproducibility, and difficulties with explainability in high-dimensional recommendation settings. We present a Decision Transformer (DT) based framework that reframes policy learning as return-conditioned supervised learning, improving robustness, scalability, and modeling flexibility. Our contributions include a real-world comparison with CQL, a multi-reward design suitable for non-episodic tasks, a quantile regression approach to return-to-go conditioning, and a production-ready system with circular buffer-based sequence processing for near-real-time inference. Extensive offline and online experiments in a deployed notification system show that our approach improves notification utility and overall session activity while minimizing user fatigue. Compared to a multi-objective CQL-based agent, the DT-based approach achieved a +0.72% increase in sessions for notification decision-making at LinkedIn by making notification recommendation more relevant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。