arXiv:2412.00301cs.LGcs.GT2024-12被引 1

兼顾双方福利的匹配市场学习算法,提升整体效率与公平性

Bandit Learning in Matching Markets: Utilitarian and Rawlsian Perspectives

  • 采用分阶段探索-利用策略学习未知偏好
  • 在稳定匹配下实现双侧福利优化,提升整体收益与公平性
  • 适合需兼顾效率与公平的实时匹配场景

双边匹配市场在学区分配、住院医师匹配、电动车充电、网约车调度和推荐系统中应用广泛。传统模型假设偏好已知,但在现代市场中,偏好往往未知,需通过学习获得。例如在线招聘中,公司难以预先知晓对所有候选人的偏好。近期研究将匹配市场建模为多臂老虎机问题,但通常只优化单方利益,导致另一方结果恶化。本文从福利主义视角出发,同时考虑两种指标:(1)功利主义福利,(2)罗尔斯主义福利,并保持市场稳定性。针对这两种度量,我们提出了基于周期性探索-利用(ETC)的算法,并分析其遗憾界。最后通过模拟实验评估了福利表现与市场稳定性。

原文摘要 · Abstract (English)

Two-sided matching markets have demonstrated significant impact in many real-world applications, including school choice, medical residency placement, electric vehicle charging, ride sharing, and recommender systems. However, traditional models often assume that preferences are known, which is not always the case in modern markets, where preferences are unknown and must be learned. For example, a company may not know its preference over all job applicants a priori in online markets. Recent research has modeled matching markets as multi-armed bandit (MAB) problem and primarily focused on optimizing matching for one side of the market, while often resulting in a pessimal solution for the other side. In this paper, we adopt a welfarist approach for both sides of the market, focusing on two metrics: (1) Utilitarian welfare and (2) Rawlsian welfare, while maintaining market stability. For these metrics, we propose algorithms based on epoch Explore-Then-Commit (ETC) and analyze their regret bounds. Finally, we conduct simulated experiments to evaluate both welfare and market stability.

匹配市场强化学习福利优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。