arXiv:2410.17086cs.GTcs.LG2024-10被引 12

用信息优势引导自私个体主动探索,平衡集体收益与个人动机。

Exploration and Persuasion

  • 通过不完全披露信息,利用信息差引导个体选择探索
  • 在多臂赌博机框架下实现探索与利用的最优平衡
  • 适合研究激励机制、博弈学习或强化学习的读者

如何激励自利主体在偏好利用的情况下仍愿意探索?考虑一群在不确定性中做决策的自利主体,他们通过探索获取新信息,再利用这些信息做出更好决策。虽然集体需要在探索与利用间取得平衡,但个体激励却偏向于利用——因为探索成本由个体承担,而收益由未来众多主体共享。本文提出“激励探索”机制:由一位能观察决策结果且可沟通但无强制力的“委托人”进行策略性推荐。关键在于信息不对称:委托人从多方收集信息,比任何单一主体更了解全局。成功的关键是委托人不完全披露知识。该问题融合了机器学习中的多臂赌博机与理论经济学中的贝叶斯说服。若所有主体都听从建议,委托人面临标准多臂赌博机问题;与单个主体互动则等价于贝叶斯说服。本文以该特殊情形为例,将多臂赌博机子问题转化为说服框架求解。

原文摘要 · Abstract (English)

How to incentivize self-interested agents to explore when they prefer to exploit? Consider a population of self-interested agents that make decisions under uncertainty. They "explore" to acquire new information and "exploit" this information to make good decisions. Collectively they need to balance these two objectives, but their incentives are skewed toward exploitation. This is because exploration is costly, but its benefits are spread over many agents in the future. "Incentivized Exploration" addresses this issue via strategic communication. Consider a benign ``principal" which can communicate with the agents and make recommendations, but cannot force the agents to comply. Moreover, suppose the principal can observe the agents' decisions and the outcomes of these decisions. The goal is to design a communication and recommendation policy which (i) achieves a desirable balance between exploration and exploitation, and (ii) incentivizes the agents to follow recommendations. What makes it feasible is "information asymmetry": the principal knows more than any one agent, as it collects information from many. It is essential that the principal does not fully reveal all its knowledge to the agents. Incentivized exploration combines two important problems in, resp., machine learning and theoretical economics. First, if agents always follow recommendations, the principal faces a multi-armed bandit problem: essentially, design an algorithm that balances exploration and exploitation. Second, interaction with a single agent corresponds to "Bayesian persuasion", where a principal leverages information asymmetry to convince an agent to take a particular action. We provide a brief but self-contained introduction to each problem through the lens of incentivized exploration, solving a key special case of the former as a sub-problem of the latter.

激励机制多臂赌博机贝叶斯说服

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。