用多尺度强化学习让AI系统兼顾短期互动与长期用户留存
MultiScale Contextual Bandits for Long Term Objectives
- 构建多尺度策略框架,用短期数据为长期目标提供层次化先验
- 在推荐与对话系统上实现比传统方法更优的长期用户留存率
- 适合关注长期用户体验的AI系统设计者与研究者
AI系统(如推荐系统、聊天机器人)从用户交互中获取的反馈是关键训练数据。尽管短期反馈(如点击、参与度)广泛用于训练,但优化短期反馈并不一定带来理想的长期目标。直接优化长期目标具有挑战性,我们识别出短期干预(如排序)与长期反馈(如用户留存)之间的时间尺度不匹配是主要障碍。为此,我们提出多尺度策略学习框架,使系统能在多个相互依赖的时间尺度上行动并优化反馈。基于PAC-Bayes理论,我们证明低时间尺度上丰富的数据可作为高时间尺度上的数据相关层次先验,从而加速稀缺数据场景下的学习。结果表明,各层级策略均有效优化长期目标。我们在推荐和对话系统三个任务中实例化该框架,提出多尺度离线策略强化学习(MSBL),验证了其有效性。
原文摘要 · Abstract (English)
The feedback that AI systems (e.g., recommender systems, chatbots) collect from user interactions is a crucial source of training data. While short-term feedback (e.g., clicks, engagement) is widely used for training, there is ample evidence that optimizing short-term feedback does not necessarily achieve the desired long-term objectives. Unfortunately, directly optimizing for long-term objectives is challenging, and we identify the disconnect in the timescales of short-term interventions (e.g., rankings) and the long-term feedback (e.g., user retention) as one of the key obstacles. To overcome this disconnect, we introduce the framework of MultiScale Policy Learning to contextually reconcile that AI systems need to act and optimize feedback at multiple interdependent timescales. Following a PAC-Bayes motivation, we show how the lower timescales with more plentiful data can provide a data-dependent hierarchical prior for faster learning at higher scales, where data is more scarce. As a result, the policies at all levels effectively optimize for the long-term. We instantiate the framework with MultiScale Off-Policy Bandit Learning (MSBL) and demonstrate its effectiveness on three tasks relating to recommender and conversational systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。