考虑人类信任度的多臂赌博机算法,提升实际应用中的决策效果。
Minimax-optimal trust-aware multi-armed bandits
- 引入动态信任模型,让推荐策略随信任变化调整
- 证明传统算法在信任缺失时性能下降,新算法可逼近最优后悔界
- 适合关注人机协作中信任机制的系统设计者
多臂赌博机(MAB)算法在顺序决策中表现优异,但其前提假设是人类能完全执行推荐策略。然而,现有方法常忽略人类对学习算法的信任问题:当信任不足时,人类可能偏离推荐策略,导致学习性能下降。为此,本文将动态信任模型融入标准MAB框架,假设推荐策略与实际执行策略受信任水平影响,而信任又随推荐质量动态演化。我们推导了存在信任问题下的最小最大后悔界,证明了经典算法如上置信界(UCB)存在次优性。为此,提出一种两阶段信任感知算法,理论证明可达到近似最优统计性能。模拟实验验证了该算法在处理信任问题上的显著优势。
原文摘要 · Abstract (English)
Multi-armed bandit (MAB) algorithms have achieved significant success in sequential decision-making applications, under the premise that humans perfectly implement the recommended policy. However, existing methods often overlook the crucial factor of human trust in learning algorithms. When trust is lacking, humans may deviate from the recommended policy, leading to undesired learning performance. Motivated by this gap, we study the trust-aware MAB problem by integrating a dynamic trust model into the standard MAB framework. Specifically, it assumes that the recommended and actually implemented policy differs depending on human trust, which in turn evolves with the quality of the recommended policy. We establish the minimax regret in the presence of the trust issue and demonstrate the suboptimality of vanilla MAB algorithms such as the upper confidence bound (UCB) algorithm. To overcome this limitation, we introduce a novel two-stage trust-aware procedure that provably attains near-optimal statistical guarantees. A simulation study is conducted to illustrate the benefits of our proposed algorithm when dealing with the trust issue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。