arXiv:2601.03423cs.CLcs.AI2026-01被引 1

不用重训就能让新大模型用上老临床模型的知识

Training-Free Adaptation of New-Generation LLMs using Legacy Clinical Models

  • 用对比解码从旧临床模型中提取医疗信号注入新模型
  • 在6个任务上平均比现有方法高17.6%至41.4%
  • 适合算力不足却想用最新大模型的医疗机构

通过持续预训练和指令微调适应临床领域需为每代新模型重新训练,成本高昂。我们提出跨架构代理调优(CAPT),一种无需训练的模型集成方法,可利用现有临床模型实现对最先进通用领域模型的适配。CAPT支持词汇表不重叠的模型,通过对比解码选择性注入临床相关信号,同时保持通用模型的推理能力和语言流畅性。在六个临床分类与文本生成任务中,使用新一代通用模型与旧一代临床模型的CAPT方案,始终优于单独使用任一模型或现有最优集成方法(平均比UniTE高出17.6%,比代理调优高出41.4%)。通过词级分析与医生案例研究,我们证实CAPT增强了临床可操作语言,减少了上下文错误,提升了临床特异性。该技术特别适合无法承担迭代临床训练、但希望采用新兴通用模型进展的医疗机构。

原文摘要 · Abstract (English)

Adapting language models to the clinical domain through continued pretraining and instruction tuning requires costly retraining for each new model generation. We propose Cross-Architecture Proxy Tuning (CAPT), a model-ensembling approach that enables training-free adaptation of state-of-the-art general-domain models using existing clinical models. CAPT supports models with disjoint vocabularies, leveraging contrastive decoding to selectively inject clinically relevant signals while preserving the general-domain model's reasoning and fluency. On six clinical classification and text-generation tasks, CAPT with a new-generation general-domain model and an older-generation clinical model consistently outperforms both models individually and state-of-the-art ensembling approaches (average +17.6\% over UniTE, +41.4\% over proxy tuning across tasks). Through token-level analysis and physician case studies, we demonstrate that CAPT amplifies clinically actionable language, reduces context errors, and increases clinical specificity. This technique especially benefits healthcare institutions with constrained computational capacity that cannot support iterative clinical training and want to adopt emerging general-domain model advances.

大模型适配医疗AI零训练模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。