考虑数据支持度的推荐系统离线策略选择方法
CASP: Support-Aware Offline Policy Selection for Two-Stage Recommender Systems
- 引入支持度惩罚项,结合双重稳健估计算法
- 在真实数据与模拟中均选出了更可靠策略
- 适合需要稳定推荐效果的工业级系统
两阶段推荐系统先生成候选集,再对候选集进行排序。由于生成器决定了排序器可用的项目,更换生成器会同时改变策略价值和用于估计该价值的数据支持。这导致了标准单阶段目标无法捕捉的离线选择问题:一个策略可能在检索得分或原始离线策略值估计上表现良好,但仍不可靠,若其依赖于支持度弱的生成器-项目组合。我们提出CAS(Coupled Action-Set Pessimism),一种针对有限策略库的两阶段推荐策略的显式支持感知离线选择器。CAS将双重稳健值估计与支持负担惩罚相结合。我们证明了忽略下游延续值的阶段规则可能任意次优,并推导出针对保守选择的总体、有限类及重构倾向性保证。在模拟实验和重构的MovieLens 1M应用中,当估计值与支持可信度存在矛盾时,CAS能选出负担更低的策略。
原文摘要 · Abstract (English)
Two-stage recommender systems first choose a candidate generator and then rank items within the generated set. Because the generator decides which items are available to the ranker, changing the generator changes both the policy value and the data support used to estimate that value. This creates an offline selection problem that standard single-stage objectives do not capture: a policy may look good under a retrieval score or a raw off-policy value estimate, but still be unreliable if it depends on weakly supported generator-item pairs. We propose CASP (Coupled Action-Set Pessimism), a support-aware offline selector for finite libraries of two-stage recommender policies. CASP combines doubly robust value estimation with a support-burden penalty. We show that stagewise rules that ignore downstream continuation value can be arbitrarily suboptimal, and we derive population, finite-class, and reconstructed-propensity guarantees for conservative selection. In simulations and a reconstructed MovieLens 1M application, CASP selects lower-burden policies when estimated value and support credibility are in tension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。