arXiv:2608.11560cs.LG2026-08中稿 · the 5th Workshop o…

提出诊断协议,解决离线评估误导问题,确保推荐系统真正有效。

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

  • 设计分步诊断流程,检验奖励与策略的对齐性和可学习性。
  • 实验证明:静态评估常误判,动态学习后效果差异显著扩大。
  • 适合需在延迟反馈下做策略选择的个性化推荐场景。

用上下文多臂老虎机(CMAB)个性化营销信息能带来实际商业价值,但最终衡量指标——转化率——需数周后才可获得,无法用于在线学习。团队通常用快速代理奖励训练带权算法,并需判断是否值得采用复杂模型而非发送单一最优信息。然而,仅依赖传统离线评估方法(批量离线策略估计、边际臂区分测试、置信区间)在延迟反馈下会系统性误导。本文提出有序诊断协议,从两个维度筛选候选奖励与策略:对齐性(优化奖励是否推动核心目标)和可学习性(带权算法能否识别最优策略)。在公开基准和可控合成生成器中验证该协议有效性,并应用于大型电商平台推送系统(5个选项,1个分组)。发现两大规律:(N1) 单一离线数值可能误排序奖励;密度更高的奖励信号使算法学习能力更强,静态评估中看似持平的奖励在在线学习后差距拉大。(N2) 若无法提前确定最优单条信息,则基于用户的策略本质上是避免选错,体现为鲁棒性而非真实个性化,“个性化溢价”易被夸大。本研究贡献为方法论层面:有序诊断协议、两条关键洞见及全流程应用经验。

原文摘要 · Abstract (English)

Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message. Settling both decisions with the usual offline checks - a batch off-policy estimate, a marginal arm-discrimination test, a confidence interval - can mislead systematically under delayed feedback. We give an ordered diagnostic protocol that screens a reward-and-policy candidate on two axes, alignment (does optimizing the reward move the north-star?) and learnability (can the bandit identify the reward-optimal policy?), before trusting any reported lift. We validate it where the truth is known - a public off-policy-evaluation benchmark and a controllable synthetic generator - and illustrate it on a deployed large-marketplace push system (where, with five arms and one split, the evidence is directional rather than powered). Two lessons recur. (N1) A single offline number can mis-rank rewards: a denser reward signal gives the bandit more to learn from, so rewards that look tied in a static estimate pull apart once learning happens online. (N2) If you cannot tell in advance which single message is best, a per-user policy partly just avoids betting on the wrong one - that looks like personalization but is really robustness, so a "personalization premium" is easily overstated. Our contribution is methodological rather than algorithmic: the ordered protocol, the two lessons it surfaces, and the end-to-end experience of applying it to a delayed-feedback CMAB.

推荐系统离线评估延迟反馈决策诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。