arXiv:2602.10608stat.MLcs.LG2026-02JMLR被引 1

用经验似然做贝叶斯推断,精准评估多策略在小样本下的表现

Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood

  • 基于经验似然构建贝叶斯框架,实现多策略联合推断
  • 小样本下仍能准确估计策略价值的不确定性
  • 适合需要可靠比较与置信度的政策分析场景

策略推断在上下文老虎机问题中至关重要。本文提出一种基于经验似然的贝叶斯推断方法,用于在有限样本条件下联合分析多个上下文老虎机策略。该方法对小样本具有鲁棒性,可为策略价值评估提供精确的不确定性量化。此外,它支持灵活的策略比较推断,并实现完整的不确定性分析。通过蒙特卡洛模拟和青少年身体质量指数数据集的应用验证了该方法的有效性。

原文摘要 · Abstract (English)

Policy inference plays an essential role in the contextual bandit problem. In this paper, we use empirical likelihood to develop a Bayesian inference method for the joint analysis of multiple contextual bandit policies in finite sample regimes. The proposed inference method is robust to small sample sizes and is able to provide accurate uncertainty measurements for policy value evaluation. In addition, it allows for flexible inferences on policy comparison with full uncertainty quantification. We demonstrate the effectiveness of the proposed inference method using Monte Carlo simulations and its application to an adolescent body mass index data set.

贝叶斯推断上下文老虎机不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。