系统梳理强化学习中预训练模型复用的实证研究,揭示有效场景与方法局限。
Combining Trained Models in Reinforcement Learning
- 基于PRISMA框架整合15项实证研究,分析迁移效果影响因素
- 发现结构相似任务或带对齐机制时复用效果更优
- 多数结果受限于特定场景,对比实验常缺乏算力公平性
深度强化学习在Atari和围棋等任务中表现优异,但仍面临样本成本高、泛化能力弱的问题。为缓解此问题,研究者常通过迁移、知识蒸馏、集成或联邦训练等方式复用预训练模型,而非从零开始训练。然而相关研究分散,不同任务、基线与计算预算差异导致比较困难。本文采用PRISMA指南,系统回顾了来自IEEE Xplore、ACM Digital Library及引文追踪的589条记录,筛选出570条唯一文献,评估89篇全文,最终纳入15项实证研究进行主合成分析。从源-目标任务相似性、复用模型多样性、与从头训练基线的公平性三方面进行定性分析。结果显示:第一,在源任务与目标任务共享显著结构或包含显式门控/对齐机制时,复用效果更佳;第二,集成与联邦聚合的证据虽有前景但稀少,且多局限于狭窄场景;第三,多数研究未进行算力匹配对比,削弱了其关于效率优势的主张。本文贡献包括缩小并统一的综述范围、研究层面的实证证据整合,以及一个有待验证的‘独立性谱系’假设,建议未来基准测试以此为参考。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) has delivered strong results in domains such as Atari and Go, but it still suffers from high sample cost and weak transfer beyond the training setting. A common response is to reuse information from previously trained models through transfer, distillation, ensemble methods, or federated training instead of learning each target task from random initialization. The literature on these mechanisms is fragmented, and published comparisons are hard to interpret because tasks, baselines, and compute budgets differ. This paper presents a PRISMA-guided systematic review of empirical studies on pretrained knowledge reuse in DRL. Starting from 589 records retrieved from IEEE Xplore, the ACM Digital Library, and citation tracing, we screened 570 unique records and assessed 89 full texts. After applying the final eligibility criteria, 15 empirical studies remained in the main synthesis. We analyzed them qualitatively across three factors: source-target similarity, diversity among reused models, and the fairness of comparisons against from-scratch baselines. Three patterns recur across the surviving corpus. First, positive results are concentrated in settings where source and target tasks share substantial structure or where the method includes an explicit gating or alignment mechanism. Second, evidence for ensembles and federated aggregation is promising but sparse and mostly limited to narrow settings. Third, compute-matched comparisons are rare, which weakens claims about efficiency gains over stronger single-agent baselines. The paper contributes a narrower and internally consistent review scope, a study-level synthesis of empirical evidence, and a provisional independence spectrum that should be treated as a hypothesis for future benchmarking rather than a validated metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。