在信息不对称下实现高效知识迁移,提升强化学习的样本效率。
The Sample Complexity of Online Strategic Decision Making with Information Asymmetry and Knowledge Transportability
- 设计非独立同分布动作策略,识别隐藏混淆因子
- 理论证明达到ε最优策略的样本复杂度为O(1/ε²)
- 适用于需跨环境迁移知识的在线博弈场景
信息不对称是多智能体系统中的普遍现象,尤其在经济与社会科学中表现显著。各智能体基于私有信息制定策略以最大化自身收益,这种策略行为常因混杂变量引入复杂性。同时,目标环境实验困难导致知识迁移面临挑战,需从数据丰富的环境转移知识。本文探讨在线学习中的核心问题:能否在要求知识迁移的前提下,利用非独立同分布的动作来学习混杂因素?我们提出一种样本高效的算法,在在线战略交互模型框架下,能准确识别系统动态,并有效应对知识迁移难题。该方法在理论上保证了学习到ε-最优策略,样本复杂度为O(1/ε²),具有紧致性。
原文摘要 · Abstract (English)
Information asymmetry is a pervasive feature of multi-agent systems, especially evident in economics and social sciences. In these settings, agents tailor their actions based on private information to maximize their rewards. These strategic behaviors often introduce complexities due to confounding variables. Simultaneously, knowledge transportability poses another significant challenge, arising from the difficulties of conducting experiments in target environments. It requires transferring knowledge from environments where empirical data is more readily available. Against these backdrops, this paper explores a fundamental question in online learning: Can we employ non-i.i.d. actions to learn about confounders even when requiring knowledge transfer? We present a sample-efficient algorithm designed to accurately identify system dynamics under information asymmetry and to navigate the challenges of knowledge transfer effectively in reinforcement learning, framed within an online strategic interaction model. Our method provably achieves learning of an $ε$-optimal policy with a tight sample complexity of $O(1/ε^2)$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。