用元学习提升重复博弈中的说服效率,理论证明更优收敛速度。
Meta-Learning for Repeated Bayesian Persuasion
- 设计元说服算法,利用多轮博弈间相似性加速学习
- 在全反馈与弱反馈下均实现更优的后悔率,优于现有最优结果
- 既适用于结构化任务,也兼容任意游戏序列,适合动态说服场景
经典贝叶斯说服研究单一战略互动中发送方如何通过精心设计的信号策略影响接收方。但在许多现实环境中,此类互动会反复发生于多个博弈中,为利用任务间的结构相似性创造了机会。本文提出元说服(Meta-Persuasion)算法,在在线贝叶斯说服(OBP)和马尔可夫说服过程(MPP)框架下,首次建立了全反馈与带奖赏反馈设置的理论结果。我们证明所提元说服算法在自然的任务相似性假设下,实现了更优的后悔率,改进了现有最优的收敛速率。同时,当博弈序列任意选取时,仍能恢复标准单次博弈的保证。最后,数值实验验证了后悔率的提升及元学习在重复说服环境中的优势。
原文摘要 · Abstract (English)
Classical Bayesian persuasion studies how a sender influences receivers through carefully designed signaling policies within a single strategic interaction. In many real-world environments, such interactions are repeated across multiple games, creating opportunities to exploit structural similarity across tasks. In this work, we introduce Meta-Persuasion algorithms, establishing the first line of theoretical results for both full-feedback and bandit-feedback settings in the Online Bayesian Persuasion (OBP) and Markov Persuasion Process (MPP) frameworks. We show that our proposed meta-persuasion algorithms achieve provably sharper regret rates under natural notions of task similarity, improving upon the best-known convergence rates for both OBP and MPP. At the same time, they recover the standard single-game guarantees when the sequence of games is picked arbitrarily. Finally, we complement our theoretical analysis with numerical experiments that highlight our regret improvements and the benefits of meta-learning in repeated persuasion environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。