arXiv:2411.19395stat.MLcs.AI2024-11被引 1

用可解释概念降低离线策略评估方差,提升结果可信度。

Concept-driven Off Policy Evaluation

  • 引入可解释概念构建新型离线评估方法,从源头减少方差。
  • 实验表明,该方法在合成与真实数据上均显著提升评估性能。
  • 结果可解释,支持针对特定概念进行干预优化。

使用批量数据评估离线策略面临样本量有限导致的高方差问题。为改善离线策略评估(OPE),需识别并缓解方差来源。近期研究表明,利用人类可解释的概念能提升预测效果并增强理解。本文提出将概念融入OPE以降低方差。我们引入一类基于概念的OPE估计器,证明当概念已知且预定义时,这些估计器保持无偏性并有效降低方差。由于现实应用中常缺乏预定义概念,我们进一步设计端到端算法,自动学习可解释、简洁且多样化的参数化概念,专为方差最小化优化。在合成与真实数据集上的实验表明,已知与学习得到的概念型估计器均显著提升OPE性能。关键的是,与其它方法不同,此类估计器具备良好可解释性,支持对特定概念进行针对性干预,从而进一步提升评估质量。

原文摘要 · Abstract (English)

Evaluating off-policy decisions using batch data poses significant challenges due to limited sample sizes leading to high variance. To improve Off-Policy Evaluation (OPE), we must identify and address the sources of this variance. Recent research on Concept Bottleneck Models (CBMs) shows that using human-explainable concepts can improve predictions and provide better understanding. We propose incorporating concepts into OPE to reduce variance. Our work introduces a family of concept-based OPE estimators, proving that they remain unbiased and reduce variance when concepts are known and predefined. Since real-world applications often lack predefined concepts, we further develop an end-to-end algorithm to learn interpretable, concise, and diverse parameterized concepts optimized for variance reduction. Our experiments with synthetic and real-world datasets show that both known and learned concept-based estimators significantly improve OPE performance. Crucially, we show that, unlike other OPE methods, concept-based estimators are easily interpretable and allow for targeted interventions on specific concepts, further enhancing the quality of these estimators.

离线评估可解释性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。