arXiv:2504.12086cs.LG2025-04被引 1

解决推荐系统中奖励延迟反馈的神经上下文老虎机问题

Neural Contextual Bandits Under Delayed Feedback Constraints

  • 基于上置信界策略,设计新算法处理随机延迟的反馈
  • 理论证明在T期内累积悔悟有上界,性能可保证
  • 适合在线推荐、临床试验等延迟反馈场景

本文提出一种新算法——延迟神经UCB(Delayed NeuralUCB),用于解决神经上下文老虎机中奖励反馈延迟的问题。该问题常见于在线推荐系统和临床试验中,因用户行为结果需时间显现而导致奖励延迟。算法采用基于上置信界(UCB)的探索策略,在奖励延迟独立同分布且服从次指数分布的假设下,推导出在长度为T的时间窗口内的累积悔悟上界。进一步提出了基于Thompson Sampling的变体算法——延迟神经TS(Delayed NeuralTS)。在真实数据集(如MNIST、Mushroom)上的数值实验表明,所提算法能有效应对不同延迟情况,优于基准方法,适用于复杂现实场景。

原文摘要 · Abstract (English)

This paper presents a new algorithm for neural contextual bandits (CBs) that addresses the challenge of delayed reward feedback, where the reward for a chosen action is revealed after a random, unknown delay. This scenario is common in applications such as online recommendation systems and clinical trials, where reward feedback is delayed because the outcomes or results of a user's actions (such as recommendations or treatment responses) take time to manifest and be measured. The proposed algorithm, called Delayed NeuralUCB, uses an upper confidence bound (UCB)-based exploration strategy. Under the assumption of independent and identically distributed sub-exponential reward delays, we derive an upper bound on the cumulative regret over a T-length horizon. We further consider a variant of the algorithm, called Delayed NeuralTS, that uses Thompson Sampling-based exploration. Numerical experiments on real-world datasets, such as MNIST and Mushroom, along with comparisons to benchmark approaches, demonstrate that the proposed algorithms effectively manage varying delays and are well-suited for complex real-world scenarios.

上下文老虎机延迟反馈神经网络推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。