arXiv:2409.02725cs.CL2024-09被引 3

用期刊影响因子筛选医学预训练数据效果不佳,少样本预训练也能保持性能。

Pre-training data selection for biomedical domain adaptation using journal impact metrics

  • 用期刊影响因子筛选PubMed数据集进行预训练
  • 减少摘要数量但保持训练步数,模型性能未下降
  • 适合关注医学NLP数据效率的科研人员

领域适应是自然语言处理中提升特定领域语言模型性能的常用方法,尤其在医学领域应用广泛。本文以PubMed为文本语料库,探讨是否可通过科学论文的质量指标(如期刊影响因子)优化预训练数据集,从而提升模型性能。实验采用两个简单的期刊影响度量指标,对BERT模型持续在不同子集的PubMed数据上进行预训练,并在BLURB基准的医学语言理解任务上评估结果。结果显示,基于影响因子的数据筛选并未带来性能提升;但发现,在保持相同训练步数的前提下,使用更少的摘要进行预训练,模型性能并未明显下降。

原文摘要 · Abstract (English)

Domain adaptation is a widely used method in natural language processing (NLP) to improve the performance of a language model within a specific domain. This method is particularly common in the biomedical domain, which sees regular publication of numerous scientific articles. PubMed, a significant corpus of text, is frequently used in the biomedical domain. The primary objective of this study is to explore whether refining a pre-training dataset using specific quality metrics for scientific papers can enhance the performance of the resulting model. To accomplish this, we employ two straightforward journal impact metrics and conduct experiments by continually pre-training BERT on various subsets of the complete PubMed training set, we then evaluate the resulting models on biomedical language understanding tasks from the BLURB benchmark. Our results show that pruning using journal impact metrics is not efficient. But we also show that pre-training using fewer abstracts (but with the same number of training steps) does not necessarily decrease the resulting model's performance.

医学NLP数据筛选预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。