arXiv:2603.16738cs.AI2026-03被引 3

医学大模型持续学习有新基准,可测稳定与效率的权衡

MedCL-Bench: Benchmarking stability-efficiency trade-offs and scaling in biomedical continual learning

  • 构建十数据集跨任务基准,统一评估持续学习策略
  • 参数隔离法在单位算力下保留能力最强,重放法代价高但保护好
  • 多标签分类最易遗忘,适合做模型更新前的风险测试

医学语言模型需随证据和术语演进持续更新,但顺序更新易引发灾难性遗忘。尽管生物医学NLP已有诸多静态基准,却缺乏统一、任务多样且遵循标准化协议、具备任务顺序鲁棒性及算力感知报告的持续学习评估框架。本文提出MedCL-Bench,包含十个覆盖五类任务的生物医学NLP数据集,评估十一种持续学习策略在八种任务顺序下的表现,报告保留率、迁移能力与GPU小时成本。结果表明,直接顺序微调会导致先前任务性能下降;不同方法在保留能力与计算成本间形成不同前沿:参数隔离在单位GPU小时下保留最优,重放提供强保护但开销大,正则化收益有限。遗忘程度具有任务依赖性,多标签主题分类最脆弱,受限输出任务更稳健。MedCL-Bench为部署前审计模型更新提供了可复现的评估框架。

原文摘要 · Abstract (English)

Medical language models must be updated as evidence and terminology evolve, yet sequential updating can trigger catastrophic forgetting. Although biomedical NLP has many static benchmarks, no unified, task-diverse benchmark exists for evaluating continual learning under standardized protocols, robustness to task order and compute-aware reporting. We introduce MedCL-Bench, which streams ten biomedical NLP datasets spanning five task families and evaluates eleven continual learning strategies across eight task orders, reporting retention, transfer, and GPU-hour cost. Across backbones and task orders, direct sequential fine-tuning on incoming tasks induces catastrophic forgetting, causing update-induced performance regressions on prior tasks. Continual learning methods occupy distinct retention-compute frontiers: parameter-isolation provides the best retention per GPU-hour, replay offers strong protection at higher cost, and regularization yields limited benefit. Forgetting is task-dependent, with multi-label topic classification most vulnerable and constrained-output tasks more robust. MedCL-Bench provides a reproducible framework for auditing model updates before deployment.

持续学习医学AI基准测试灾难性遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。