arXiv:2607.22012cs.LG2026-07ICLR被引 1

跨域数据提升小样本下的策略评估与学习效果

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

论文配图:Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits
图 1 · 摘自论文原文
  • 利用目标域与源域历史数据联合建模,缓解数据稀缺问题
  • 在少样本、确定性采样等挑战场景下显著降低评估方差
  • 适用于医疗、推荐等需安全试错的现实场景

上下文老虎机中的离线策略评估与学习(OPE/L)因其可仅用历史日志数据安全评估和训练新策略而迅速应用于真实系统。然而,现有方法难以应对少样本数据、确定性日志策略及新增动作等常见但严峻的挑战。在个性化医疗、内容推荐、教育和广告等领域,这些情况普遍存在。现有方法因方差过大或日志数据探索不足而无法有效评估与优化。为此,我们提出跨域离线策略评估与学习的新范式:除目标域日志外,还引入其他域的历史数据。该设定广泛适用,因常可获取不同医院、国家、设备或用户群体的过往数据。我们开发了一种新估计器与策略梯度方法,联合利用目标域与源域数据,在实验中显著提升了以往无法解决场景下的评估与学习性能。

原文摘要 · Abstract (English)

Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handle many challenging but prevalent scenarios such as few-shot data, deterministic logging policies, and new actions. In many applications, such as personalized medicine, content recommendations, education, and advertising, we need to evaluate and learn new policies in the presence of these challenges. Existing methods cannot evaluate and optimize effectively in these situations due to the notorious variance issue or limited exploration in the logged data. To enable OPE/L even under these unsolved challenges, we propose a new problem setup of Cross-Domain OPE/L, where we have access not only to the logged data from the target domain in which the new policy will be implemented but also to logged datasets collected from other domains. This novel formulation is widely applicable because we can often use historical data not only from the target hospital, country, device, or user segment but also from other hospitals, countries, devices, or segments. We develop a new estimator and policy gradient method to solve OPE/L by leveraging both target and source datasets, resulting in substantially enhanced OPE/L in the previously unsolved situations in our empirical evaluations.

离线学习策略评估跨域学习上下文老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。