arXiv:2601.05383cs.LGmath.OC2026-01

用专家分类框架提升不确定环境下组合优化的模仿学习效果

Imitation Learning for Combinatorial Optimisation under Uncertainty

  • 提出三维度专家分类法:不确定性处理、最优性水平、交互方式
  • 随机专家比确定性专家学出的策略更优,互动学习减少所需示范数
  • 适用于医疗资源分配等动态决策场景,适合做强化学习前导研究

模仿学习(IL)为大规模组合优化问题提供数据驱动的近似策略求解框架,此类问题常被建模为序列决策问题(SDPs),而精确求解方法计算上不可行。当前研究中,生成训练示范的‘专家’角色尚未系统化,其建模假设、计算特性与学习性能影响各异。本文构建了不确定性下组合优化模仿学习的专家分类体系,从三个核心维度划分:(i)不确定性处理方式;(ii)最优性水平,区分任务最优与近似专家;(iii)与学习者的交互模式,涵盖单次监督到迭代互动。进一步识别其他相关特征类别。基于此,提出通用数据聚合(DAgger)框架,支持多轮专家查询、专家聚合与灵活交互策略。在具有随机到达与容量约束的动态医生-患者分配问题上进行实验。结果表明,由随机专家训练的策略持续优于确定性或全信息专家;互动学习在减少专家示范数量的同时提升解质量。当随机优化计算困难时,聚合的确定性专家仍具有效性。

原文摘要 · Abstract (English)

Imitation learning (IL) provides a data-driven framework for approximating policies for large-scale combinatorial optimisation problems formulated as sequential decision problems (SDPs), where exact solution methods are computationally intractable. A central but underexplored aspect of IL in this context is the role of the \emph{expert} that generates training demonstrations. Existing studies employ a wide range of expert constructions, yet lack a unifying framework to characterise their modelling assumptions, computational properties, and impact on learning performance. This paper introduces a systematic taxonomy of experts for imitation learning in combinatorial optimisation under uncertainty. The literature is classified along three principal dimensions: (i) treatment of uncertainty; (ii) level of optimality, distinguishing task-optimal and approximate experts; and (iii) interaction mode with the learner, ranging from one-shot supervision to iterative, interactive schemes. We further identify additional categories capturing other relevant expert characteristics. Building on this taxonomy, we propose a generalised Dataset Aggregation (DAgger) framework that accommodates multiple expert queries, expert aggregation, and flexible interaction strategies. The proposed framework is evaluated on a dynamic physician-to-patient assignment problem with stochastic arrivals and capacity constraints. Computational experiments compare learning outcomes across expert types and interaction regimes. The results show that policies learned from stochastic experts consistently outperform those learned from deterministic or full-information experts, while interactive learning improves solution quality using fewer expert demonstrations. Aggregated deterministic experts provide an effective alternative when stochastic optimisation becomes computationally challenging.

模仿学习组合优化不确定性交互学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。