arXiv:2602.20465cs.GTcs.LG2026-02

让推荐系统在不依赖用户先验的情况下,激励各方长期合作探索。

Prior-Agnostic Incentive-Compatible Exploration

  • 用交换后悔值控制机制,使代理按推荐行动
  • 即使先验不同,也能逼近贝叶斯纳什均衡
  • 适合动态推荐平台与不确定时序的场景

在多人顺序决策的强化学习设置中,长期优化需持续探索,但短期次优动作对个体无激励。当一个长期存在的平台向一系列不同用户推荐动作时,会出现激励错配:平台愿探索,但用户不愿。已有研究在静态随机环境中假设所有参与者共享同一先验,且满足贝叶斯激励相容性。本文证明,仅靠(加权)交换后悔界即可使代理在近似贝叶斯纳什均衡下忠实跟随预测,即使在动态环境、代理间存在冲突先验信念、且机制设计者完全不知其信念的情况下亦成立。为获得此类界限,必须假设代理不仅对回报有不确定性,还对自身在服务序列中的相对位置(到达时间)存在不确定性。我们进一步给出了实现自适应与加权后悔控制的具体算法。

原文摘要 · Abstract (English)

In bandit settings, optimizing long-term regret metrics requires exploration, which corresponds to sometimes taking myopically sub-optimal actions. When a long-lived principal merely recommends actions to be executed by a sequence of different agents (as in an online recommendation platform) this provides an incentive misalignment: exploration is "worth it" for the principal but not for the agents. Prior work studies regret minimization under the constraint of Bayesian Incentive-Compatibility in a static stochastic setting with a fixed and common prior shared amongst the agents and the algorithm designer. We show that (weighted) swap regret bounds on their own suffice to cause agents to faithfully follow forecasts in an approximate Bayes Nash equilibrium, even in dynamic environments in which agents have conflicting prior beliefs and the mechanism designer has no knowledge of any agents beliefs. To obtain these bounds, it is necessary to assume that the agents have some degree of uncertainty not just about the rewards, but about their arrival time -- i.e. their relative position in the sequence of agents served by the algorithm. We instantiate our abstract bounds with concrete algorithms for guaranteeing adaptive and weighted regret in bandit settings.

强化学习激励相容多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。