用强化学习动态组合专家策略,提升在线匹配系统效率
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
- 通过优势加权融合多个专家策略,实现自适应决策
- 在器官匹配模拟中,系统效率高于单个专家和传统强化学习
- 支持大规模状态空间,兼具可解释性与高效性,适合复杂资源调度
在线匹配问题广泛存在于云服务、在线市场及器官交换网络等复杂系统中,及时且合理的决策对系统性能至关重要。传统启发式方法虽简单易懂,但通常仅适用于特定运行场景,条件变化时易导致效率下降。本文提出一种基于强化学习的方法,通过数据驱动方式动态协调一组专家策略,发挥其互补优势。基于Adv2框架(Jonckheere et al., 2024),该方法采用基于优势的权重更新机制,并可自然扩展至仅有估计值函数可用的情况。我们建立了期望与高概率下的后悔上界,并推导出时序差分学习的新型有限时间偏差界,使在固定步长和非平稳动态下仍能可靠估计优势。为支持可扩展性,引入神经网络演员-评论家架构,在保持可解释性的同时泛化至大状态空间。在随机匹配模型(包括器官交换场景)上的仿真表明,该协同策略收敛更快,系统整体效率显著优于单个专家及传统强化学习基线。结果表明,结构化自适应学习能有效提升复杂资源分配与决策过程的建模与管理能力。
原文摘要 · Abstract (English)
Online matching problems arise in many complex systems, from cloud services and online marketplaces to organ exchange networks, where timely, principled decisions are critical for maintaining high system performance. Traditional heuristics in these settings are simple and interpretable but typically tailored to specific operating regimes, which can lead to inefficiencies when conditions change. We propose a reinforcement learning (RL) approach that learns to orchestrate a set of such expert policies, leveraging their complementary strengths in a data-driven, adaptive manner. Building on the Adv2 framework (Jonckheere et al., 2024), our method combines expert decisions through advantage-based weight updates and extends naturally to settings where only estimated value functions are available. We establish both expectation and high-probability regret guarantees and derive a novel finite-time bias bound for temporal-difference learning, enabling reliable advantage estimation even under constant step size and non-stationary dynamics. To support scalability, we introduce a neural actor-critic architecture that generalizes across large state spaces while preserving interpretability. Simulations on stochastic matching models, including an organ exchange scenario, show that the orchestrated policy converges faster and yields higher system level efficiency than both individual experts and conventional RL baselines. Our results highlight how structured, adaptive learning can improve the modeling and management of complex resource allocation and decision-making processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。