临床T5模型比通用T5模型提升有限,跨领域泛化能力反而更差。
Are Clinical T5 Models Better for Clinical Text?
- 对比临床与通用T5模型在多个临床任务上的表现
- 临床T5仅微弱优于通用模型,跨域测试时性能下降
- 为临床大模型开发提供实证参考,适合关注泛化性的研究者
基于Transformer编码器/解码器架构的大语言模型(如T5)已成为监督学习任务的标准平台。为将此类技术引入临床领域,近期工作训练或调整了适用于临床数据的T5模型。然而,这些临床T5模型的评估及其与其它模型的比较仍不充分。临床T5模型是否优于经过FLAN调优的通用T5模型?它们在与训练集不同的新临床领域上是否具有更好的泛化能力?我们对这些模型在多个临床任务和领域中进行了全面评估。结果表明,临床T5模型相较于现有模型仅有微弱提升,且在不同领域的评估中表现更差。该结果为未来临床大模型的开发提供了重要依据。
原文摘要 · Abstract (English)
Large language models with a transformer-based encoder/decoder architecture, such as T5, have become standard platforms for supervised tasks. To bring these technologies to the clinical domain, recent work has trained new or adapted existing models to clinical data. However, the evaluation of these clinical T5 models and comparison to other models has been limited. Are the clinical T5 models better choices than FLAN-tuned generic T5 models? Do they generalize better to new clinical domains that differ from the training sets? We comprehensively evaluate these models across several clinical tasks and domains. We find that clinical T5 models provide marginal improvements over existing models, and perform worse when evaluated on different domains. Our results inform future choices in developing clinical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。