解决推荐系统中汤普森采样下的反事实推断难题
Counterfactual Inference under Thompson Sampling
- 推导出多种分布下汤普森采样的动作倾向性精确表达式
- 首次实现汤普森采样场景中离线评估的无偏估计
- 适用于推荐、广告等在线决策系统的因果推断
推荐系统是不确定性下的序列决策典范,需权衡探索与利用。汤普森采样通过概率化选择策略自然平衡此权衡,但其在反事实推断(如离线评估)中的应用受限,因现有估计器依赖动作倾向性,而该值在汤普森采样下难以获得。本文推导出多种参数与结果分布下动作倾向性的精确且高效可计算表达式,使离线策略评估器可在汤普森采样场景中使用。这为推荐系统无偏离线评估,以及在线广告、个性化等领域的因果推断提供了可行路径。
原文摘要 · Abstract (English)
Recommender systems exemplify sequential decision-making under uncertainty, strategically deciding what content to serve to users, to optimise a range of potential objectives. To balance the explore-exploit trade-off successfully, Thompson sampling provides a natural and widespread paradigm to probabilistically select which action to take. Questions of causal and counterfactual inference, which underpin use-cases like offline evaluation, are not straightforward to answer in these contexts. Specifically, whilst most existing estimators rely on action propensities, these are not readily available under Thompson sampling procedures. We derive exact and efficiently computable expressions for action propensities under a variety of parameter and outcome distributions, enabling the use of off-policy estimators in Thompson sampling scenarios. This opens up a range of practical use-cases where counterfactual inference is crucial, including unbiased offline evaluation of recommender systems, as well as general applications of causal inference in online advertising, personalisation, and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。