arXiv:2503.04655cs.LG2025-03ICLR被引 2

提出动态评估框架CLDyB,解决预训练模型持续学习评估的静态瓶颈问题。

CLDyB: Towards Dynamic Benchmarking for Continual Learning with Pre-trained Models

  • 基于马尔可夫决策过程构建动态评估框架,自动识别难任务与挑战性顺序。
  • 在生成的任务序列上,现有方法普遍表现不佳,揭示其泛化缺陷。
  • 适合研究持续学习鲁棒性、算法评估的学者,可提升方法验证可靠性。

预训练模型时代推动了持续学习(CL)研究,但预训练阶段的数据泄露隐患及静态基准的局限性日益凸显。为应对这一挑战,本文提出基于马尔可夫决策过程的动态评估框架CLDyB,能够自动生成具有挑战性的任务序列。通过蒙特卡洛树搜索确定算法依赖的困难任务与顺序,实现对多种先进CL方法的联合评估,发现其共同弱点。进一步对单个方法进行独立评估,揭示其优劣势。所生成的任务序列与源码已公开于https://github.com/szc12153/CLDyB。

原文摘要 · Abstract (English)

The advent of the foundation model era has sparked significant research interest in leveraging pre-trained representations for continual learning (CL), yielding a series of top-performing CL methods on standard evaluation benchmarks. Nonetheless, there are growing concerns regarding potential data contamination during the pre-training stage. Furthermore, standard evaluation benchmarks, which are typically static, fail to capture the complexities of real-world CL scenarios, resulting in saturated performance. To address these issues, we describe CL on dynamic benchmarks (CLDyB), a general computational framework based on Markov decision processes for evaluating CL methods reliably. CLDyB dynamically identifies inherently difficult and algorithm-dependent tasks for the given CL methods, and determines challenging task orders using Monte Carlo tree search. Leveraging CLDyB, we first conduct a joint evaluation of multiple state-of-the-art CL methods, leading to a set of commonly challenging and generalizable task sequences where existing CL methods tend to perform poorly. We then conduct separate evaluations of individual CL methods using CLDyB, discovering their respective strengths and weaknesses. The source code and generated task sequences are publicly accessible at https://github.com/szc12153/CLDyB.

持续学习动态评估预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。