arXiv:2509.00648cs.LGstat.ML2025-09

通过学习上下文-动作嵌入,降低离策略评估的误差。

Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits

  • 从离线数据中学习上下文-动作嵌入以最小化估计误差。
  • 在合成与真实数据集上,均显著优于基线方法的均方误差。
  • 适合需要高精度离策略评估的研究者使用。

我们研究有限动作空间下上下文随机决策问题中的离策略评估(OPE)。逆倾向得分(IPS)加权虽无偏但易因动作空间大或某些上下文-动作区域未充分探索而产生高方差。近期提出的边际化IPS(MIPS)通过利用动作嵌入缓解此问题,但其嵌入未最小化估计器的均方误差(MSE),且未考虑上下文信息。为此,本文提出上下文-动作嵌入学习的MIPS(CAEL-MIPS),基于对MIPS偏差与方差的理论分析,构建最小化MSE的目标函数,从离线数据中学习上下文-动作嵌入。在合成数据集与真实数据集上的实证研究表明,该方法在均方误差上优于现有基线。

原文摘要 · Abstract (English)

We consider off-policy evaluation (OPE) in contextual bandits with finite action space. Inverse Propensity Score (IPS) weighting is a widely used method for OPE due to its unbiased, but it suffers from significant variance when the action space is large or when some parts of the context-action space are underexplored. Recently introduced Marginalized IPS (MIPS) estimators mitigate this issue by leveraging action embeddings. However, these embeddings do not minimize the mean squared error (MSE) of the estimators and do not consider context information. To address these limitations, we introduce Context-Action Embedding Learning for MIPS, or CAEL-MIPS, which learns context-action embeddings from offline data to minimize the MSE of the MIPS estimator. Building on the theoretical analysis of bias and variance of MIPS, we present an MSE-minimizing objective for CAEL-MIPS. In the empirical studies on a synthetic dataset and a real-world dataset, we demonstrate that our estimator outperforms baselines in terms of MSE.

离策略评估上下文带宽嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。