arXiv:2410.08537cs.LGstat.ML2024-10被引 2

利用多源异构观测数据,学习能跨场景鲁棒泛化的个性化决策策略。

Robust Offline Policy Learning with Observational Data from Multiple Sources

  • 设计极小极大后悔目标,确保在多源混合分布下统一低后悔。
  • 算法结合双重稳健评估与无后悔学习,实现最优最坏情况后悔率。
  • 适合需要从多个不一致数据源中提取稳定决策规则的场景。

我们研究如何利用来自多个异构数据源的观测性带状反馈数据,学习一种能在多样目标设置中稳健泛化的个性化决策策略。为此,我们提出一种极小极大后悔优化目标,以保证在所有源分布混合情形下均具有统一的低后悔表现。我们开发了一种针对该目标的策略学习算法,结合双重稳健的离线策略评估技术与用于极小极大优化的无后悔学习算法。我们的后悔分析表明,该方法在所有源数据总量增加时,能达到最小化最坏情况混合后悔的适度衰减速率。分析、扩展及实验结果均验证了该方法在多源数据上学习鲁棒决策策略的优势。

原文摘要 · Abstract (English)

We consider the problem of using observational bandit feedback data from multiple heterogeneous data sources to learn a personalized decision policy that robustly generalizes across diverse target settings. To achieve this, we propose a minimax regret optimization objective to ensure uniformly low regret under general mixtures of the source distributions. We develop a policy learning algorithm tailored to this objective, combining doubly robust offline policy evaluation techniques and no-regret learning algorithms for minimax optimization. Our regret analysis shows that this approach achieves the minimal worst-case mixture regret up to a moderated vanishing rate of the total data across all sources. Our analysis, extensions, and experimental results demonstrate the benefits of this approach for learning robust decision policies from multiple data sources.

决策学习多源数据鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。