arXiv:2506.08438cs.LGcs.GT2025-06被引 2

让智能体诚实汇报,主理人高效学习最优协作策略。

Learning to Lead: Incentivizing Strategic Agents in the Dark

  • 用延迟机制让理性代理近似短期决策,减少欺骗动机。
  • 通过扇区测试与匹配恢复类型相关奖励函数,精准识别隐藏偏好。
  • 在激励约束下实现近最优 $ ilde{O}( ext{√}T)$ 学习误差,适合博弈学习场景。

我们研究了一个在线学习版本的广义委托-代理模型,其中主理人反复与具有私有类型、私有奖励并执行不可观测行为的战略性代理互动。该代理非短视,优化未来奖励的折扣和,可能战略性地虚报类型以操纵主理人的学习。主理人仅观察自身实际回报及代理报告的类型,目标是学习一个最小化战略后悔的最优协调机制。我们提出了首个可证明样本高效的算法。该方法包含三个核心部分:(i)延迟机制以激励近似短视的代理行为;(ii)创新的奖励角度估计框架,结合扇区测试与匹配程序,恢复依赖类型的奖励函数;(iii)悲观-乐观的线性上下文老虎机(LinUCB)算法,使主理人在满足代理激励约束的前提下高效探索。我们建立了主理人最优策略学习的近最优 $ ilde{O}( ext{√}T)$ 后悔界,其中 $ ilde{O}( ext{·})$ 省略对数因子。结果为设计一系列涉及私有类型与战略代理的鲁棒在线学习算法开辟了新路径。

原文摘要 · Abstract (English)

We study an online learning version of the generalized principal-agent model, where a principal interacts repeatedly with a strategic agent possessing private types, private rewards, and taking unobservable actions. The agent is non-myopic, optimizing a discounted sum of future rewards and may strategically misreport types to manipulate the principal's learning. The principal, observing only her own realized rewards and the agent's reported types, aims to learn an optimal coordination mechanism that minimizes strategic regret. We develop the first provably sample-efficient algorithm for this challenging setting. Our approach features a novel pipeline that combines (i) a delaying mechanism to incentivize approximately myopic agent behavior, (ii) an innovative reward angle estimation framework that uses sector tests and a matching procedure to recover type-dependent reward functions, and (iii) a pessimistic-optimistic LinUCB algorithm that enables the principal to explore efficiently while respecting the agent's incentive constraints. We establish a near optimal $\tilde{O}(\sqrt{T}) $ regret bound for learning the principal's optimal policy, where $\tilde{O}(\cdot) $ omits logarithmic factors. Our results open up new avenues for designing robust online learning algorithms for a wide range of game-theoretic settings involving private types and strategic agents.

在线学习博弈论激励机制代理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。