arXiv:2601.08120cs.LG2026-01AAAI被引 2

动态识别任务结构,提升上下文强化学习的迁移效果

Structure Detection for Contextual Reinforcement Learning

  • 根据任务结构自动选择最优迁移策略
  • 在多个基准上比之前方法平均提升12.49%
  • 适合复杂场景下的多任务强化学习研究者

上下文强化学习(CRL)旨在解决一组随上下文变量变化的上下文马尔可夫决策过程(CMDP)。传统方法如独立训练和多任务学习在计算成本或负迁移方面表现不佳。近期提出的基于模型的迁移学习(MBTL)通过有策略地选择少数任务进行训练并实现零样本迁移,已展现出有效性。然而,不同CMDP具有不同的结构特性,需匹配相应任务选择策略。本文提出结构检测MBTL(SD-MBTL),一个能动态识别CMDP泛化结构并自适应选择合适MBTL算法的通用框架。例如,在发现‘山形’结构时(即随着上下文差异增大,泛化性能下降),提出M/GP-MBTL,可动态切换高斯过程与聚类方法。在合成数据及涵盖连续控制、交通控制和农业管理的CRL基准上的大量实验表明,M/GP-MBTL在综合指标上优于最强基线方法12.49%。结果表明,在线结构检测对复杂CRL环境中的源任务选择具有重要指导意义。

原文摘要 · Abstract (English)

Contextual Reinforcement Learning (CRL) tackles the problem of solving a set of related Contextual Markov Decision Processes (CMDPs) that vary across different context variables. Traditional approaches--independent training and multi-task learning--struggle with either excessive computational costs or negative transfer. A recently proposed multi-policy approach, Model-Based Transfer Learning (MBTL), has demonstrated effectiveness by strategically selecting a few tasks to train and zero-shot transfer. However, CMDPs encompass a wide range of problems, exhibiting structural properties that vary from problem to problem. As such, different task selection strategies are suitable for different CMDPs. In this work, we introduce Structure Detection MBTL (SD-MBTL), a generic framework that dynamically identifies the underlying generalization structure of CMDP and selects an appropriate MBTL algorithm. For instance, we observe Mountain structure in which generalization performance degrades from the training performance of the target task as the context difference increases. We thus propose M/GP-MBTL, which detects the structure and adaptively switches between a Gaussian Process-based approach and a clustering-based approach. Extensive experiments on synthetic data and CRL benchmarks--covering continuous control, traffic control, and agricultural management--show that M/GP-MBTL surpasses the strongest prior method by 12.49% on the aggregated metric. These results highlight the promise of online structure detection for guiding source task selection in complex CRL environments.

强化学习迁移学习结构检测多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。