用Transformer模型提升医学文献标题的自动分类准确率
Enhancing Automatic PT Tagging for MEDLINE Citations Using Transformer-Based Models
- 采用BERT和DistilBERT等预训练模型处理文献元数据
- 多标签分类器使术语标注准确率显著提升
- 适合需要高效文献管理的研究人员与数据库维护者
我们研究了利用预训练的基于Transformer的模型BERT和DistilBERT,从MEDLINE文献元数据中预测医学主题词(MeSH)出版类型(PT)的可行性。该研究旨在解决当前自动化标引依赖过时自然语言处理算法的问题。通过评估单一多标签分类器和二分类器集成方法,提升了生物医学文献的检索效果。结果表明,Transformer模型能显著提高出版类型标注的准确性,为可扩展、高效的生物医学文献标引提供了可能。
原文摘要 · Abstract (English)
We investigated the feasibility of predicting Medical Subject Headings (MeSH) Publication Types (PTs) from MEDLINE citation metadata using pre-trained Transformer-based models BERT and DistilBERT. This study addresses limitations in the current automated indexing process, which relies on legacy NLP algorithms. We evaluated monolithic multi-label classifiers and binary classifier ensembles to enhance the retrieval of biomedical literature. Results demonstrate the potential of Transformer models to significantly improve PT tagging accuracy, paving the way for scalable, efficient biomedical indexing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。