考虑延迟反馈的资源分配新方法,提升公平性与长期效果。
Bi-Level Contextual Bandits for Individualized Resource Allocation under Delayed Feedback
- 分层设计:上层优化群体预算,下层识别最响应个体。
- 实测在教育与就业数据集上累积收益更高,适应延迟结构。
- 适合需兼顾公平、容量限制与延迟反馈的政策决策场景。
在教育、就业和医疗等高风险领域中,公平分配有限资源需平衡短期效益与长期影响,同时应对延迟结果、隐藏异质性及伦理约束。现有学习型分配框架多假设即时反馈,或忽略个体特征与干预动态的复杂交互。本文提出一种新型双层上下文老虎机框架,用于延迟反馈下的个性化资源分配,适用于动态人群、容量约束和时效性影响的现实场景。在元层级,模型优化子群体预算分配以满足公平性和操作约束;在基础层级,利用观测数据训练神经网络,识别各组内最敏感个体,同时遵守冷却期与资源特异性延迟核建模的治疗延迟效应。通过显式建模时间动态与反馈延迟,算法随新数据持续优化策略,实现更灵活自适应决策。我们在教育与劳动力发展两个真实数据集上验证该方法,结果表明其累积成效更高,更优适应延迟模式,并确保跨子群体的公平分配。研究凸显了具备延迟感知能力的数据驱动决策系统在改进制度政策与社会福利方面的潜力。
原文摘要 · Abstract (English)
Equitably allocating limited resources in high-stakes domains-such as education, employment, and healthcare-requires balancing short-term utility with long-term impact, while accounting for delayed outcomes, hidden heterogeneity, and ethical constraints. However, most learning-based allocation frameworks either assume immediate feedback or ignore the complex interplay between individual characteristics and intervention dynamics. We propose a novel bi-level contextual bandit framework for individualized resource allocation under delayed feedback, designed to operate in real-world settings with dynamic populations, capacity constraints, and time-sensitive impact. At the meta level, the model optimizes subgroup-level budget allocations to satisfy fairness and operational constraints. At the base level, it identifies the most responsive individuals within each group using a neural network trained on observational data, while respecting cooldown windows and delayed treatment effects modeled via resource-specific delay kernels. By explicitly modeling temporal dynamics and feedback delays, the algorithm continually refines its policy as new data arrive, enabling more responsive and adaptive decision-making. We validate our approach on two real-world datasets from education and workforce development, showing that it achieves higher cumulative outcomes, better adapts to delay structures, and ensures equitable distribution across subgroups. Our results highlight the potential of delay-aware, data-driven decision-making systems to improve institutional policy and social welfare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。