arXiv:2505.23720cs.LGcs.AI2025-05被引 1

提出无需金钱激励的算法,让自利卖家诚实上报商品信息

COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents

  • 设计不依赖金钱激励的机制,抑制策略性行为
  • 理论保证策略无关性与次线性后悔率
  • 适用于电商推荐等存在利益博弈的场景

本文研究多代理情境下的上下文老虎机问题,学习者需根据上下文及代理报告的臂选择最优动作以最大化系统总收益。现有工作假设代理如实报告,但在实际中如电商平台卖家可能虚报商品质量以获取推荐优势。为此,我们提出COBRA算法,在不使用任何货币激励的前提下,使代理无动机进行策略性操作,并保证激励相容性和次线性后悔率。实验结果验证了该算法在不同性能维度上的有效性。

原文摘要 · Abstract (English)

This paper considers a contextual bandit problem involving multiple agents, where a learner sequentially observes the contexts and the agent's reported arms, and then selects the arm that maximizes the system's overall reward. Existing work in contextual bandits assumes that agents truthfully report their arms, which is unrealistic in many real-life applications. For instance, consider an online platform with multiple sellers; some sellers may misrepresent product quality to gain an advantage, such as having the platform preferentially recommend their products to online users. To address this challenge, we propose an algorithm, COBRA, for contextual bandit problems involving strategic agents that disincentivize their strategic behavior without using any monetary incentives, while having incentive compatibility and a sub-linear regret guarantee. Our experimental results also validate the different performance aspects of our proposed algorithm.

强化学习机制设计上下文老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。