用行为经济学模型优化匹配市场中的智能体学习,提升抗干扰能力。
Matching Markets meet Cumulative Prospect Theory: Towards Optimal and Adversarially Robust Learning
- 引入累积前景理论建模人类非线性偏好,改进多臂赌博机学习策略。
- 在千人级选项场景下,实现与选项数无关的最优后悔率,优于现有方法。
- 设计鲁棒算法,可在奖励被恶意篡改时仍保持对数级学习误差,适合高风险场景。
我们研究了在双侧匹配市场中,基于人类中心决策模型的多智能体多臂赌博机问题。为刻画人类偏好,采用累积前景理论(CPT),通过α-霍尔德连续权重函数非线性地加权智能体行为。该理论广泛用于行为经济学和风险敏感机器学习以模拟人类决策。分析了当前最先进的学习算法在CPT扭曲奖励下的表现,获得玩家最优后悔率为 𝒪(K log T (1/Δ)^{2/α}),其中 K 为臂的数量,T 为学习周期,Δ 为玩家间最小偏好差距。注意到 Δ 的依赖关系次优,进一步通过精明选择探索阶段活跃臂集,消除了主导项中对 K 的依赖,在臂数 K 远大于玩家数 N 时达到最优后悔率。此外,考虑对抗性市场环境,即观测奖励可能被污染,提出并分析两种鲁棒算法:一种已知总污染预算,一种未知。在两种情况下均建立对数级玩家最优后悔界。
原文摘要 · Abstract (English)
We study a multi-agent multi-armed bandit problem in the competitive setup with two-sided matching markets under a human centric decision making model. To capture human preferences, we use cumulative prospect theory (CPT) that weighs the actions of the agent in a nonlinear fashion using a ($α$-Hölder continuous) weight function. CPT has been widely used in behavioral economics and risk sensitive machine learning to emulate human preferences. We analyze the state-of-the-art learning algorithm with CPT weight distorted rewards and obtain a player optimal regret of $\mathcal{O}(K\log T \left(\frac{1}Δ\right)^{2/α})$, where $K$ denotes the number of arms, $T$ is the learning horizon, and $Δ$ represents (suitably defined) players' minimum preference gap. Noticing the dependence on $Δ$ to be sub-optimal, we further improve this regret by judiciously selecting the active set of arms during exploration, which removes the dependence on $K$ in the dominant term and achieves an improved (optimal) regret guarantees in the setting where the number of arms $K$ is significantly larger than the number of players $N$. In addition, we consider adversarial markets where the observed rewards of the agents may be corrupted. We propose and analyze algorithms for robust markets with CPT as risk sensitive measure in both settings where the total corruption budget is known and where it is unknown, and establish logarithmic player-optimal regret guarantees in both cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。