arXiv:2510.22069cs.LGstat.ML2025-10被引 1

神经网络学习带异构预算的多动作强化学习策略,实现高效资源分配。

Neural Index Policies for Restless Multi-Action Bandits with Heterogeneous Budgets

  • 用神经网络学习带预算的索引值,通过可微背包层实现约束分配
  • 在数百个动作上接近最优,误差小于5%,且严格满足异构预算
  • 适合医疗、资源调度等复杂约束场景,理论与实践结合

非休息多臂老虎机(RMAB)为不确定环境下的序列决策提供可扩展框架,但经典模型假设二元动作和单一全局预算。真实场景如医疗中常涉及多种干预手段,成本和约束异质,导致传统假设失效。本文提出神经索引策略(NIP),用于具有异构预算约束的多动作RMAB。该方法利用神经网络学习带预算感知的臂-动作对索引,并通过可微背包层(建模为熵正则化最优传输问题)转化为可行分配。整个模型将索引预测与约束优化统一于端到端可微框架中,支持基于决策质量的梯度训练。网络优化目标是使诱导的占用测度逼近线性规划松弛的理论上限,连接渐近RMAB理论与实际学习。实验表明,NIP在保持异构预算约束前提下,性能接近5%的虚拟最优占用测度策略,且可扩展至数百个臂。本工作建立了一个通用、理论坚实且可扩展的学习索引策略框架,适用于复杂资源受限环境。

原文摘要 · Abstract (English)

Restless multi-armed bandits (RMABs) provide a scalable framework for sequential decision-making under uncertainty, but classical formulations assume binary actions and a single global budget. Real-world settings, such as healthcare, often involve multiple interventions with heterogeneous costs and constraints, where such assumptions break down. We introduce a Neural Index Policy (NIP) for multi-action RMABs with heterogeneous budget constraints. Our approach learns to assign budget-aware indices to arm--action pairs using a neural network, and converts them into feasible allocations via a differentiable knapsack layer formulated as an entropy-regularized optimal transport (OT) problem. The resulting model unifies index prediction and constrained optimization in a single end-to-end differentiable framework, enabling gradient-based training directly on decision quality. The network is optimized to align its induced occupancy measure with the theoretical upper bound from a linear programming relaxation, bridging asymptotic RMAB theory with practical learning. Empirically, NIP achieves near-optimal performance within 5% of the oracle occupancy-measure policy while strictly enforcing heterogeneous budgets and scaling to hundreds of arms. This work establishes a general, theoretically grounded, and scalable framework for learning index-based policies in complex resource-constrained environments.

强化学习资源分配神经网络多臂老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。