arXiv:2504.00748cs.CL2025-04

用大模型自动从文献中提取肿瘤免疫组化信息,提升研究效率。

IHC-LLMiner: Automated extraction of tumour immunohistochemical profiles from PubMed abstracts using large language models

  • 基于微调的Gemma-2模型,自动分类文献并提取免疫组化-肿瘤关联数据
  • 识别出30,481篇相关文献,对50种标记物的提取准确率达63.3%
  • 结果与公开数据高度一致,适合癌症研究者构建结构化知识库

免疫组化(IHC)在病理诊断和生物医学研究中至关重要,可揭示蛋白质表达与肿瘤生物学特征。本研究提出自动化管道IHC-LLMiner,利用大语言模型从PubMed摘要中提取IHC-肿瘤谱型。包含两个子任务:摘要分类(相关/无关)与相关摘要中的IHC-肿瘤谱型提取。最佳模型Gemma-2微调版在分类任务中达91.5%准确率与91.4 F1分数,较GPT4-O高出9.5个百分点,推理速度提升5.9倍。从初始107,759篇摘要中筛选出30,481篇相关文献。在谱型提取任务中,该模型表现最优,正确输出率达63.3%。提取的肿瘤类型与标记物经统一医学语言系统(UMLS)标准化,确保一致性,便于整体分析。结果与现有在线汇总数据高度一致,不仅补充了缺失的IHC-肿瘤组合,还提供量化评估价值。所提基于LLM的管道为大规模IHC-肿瘤谱型数据挖掘提供了实用方案,提升数据可及性与可用性,支持癌症特异性知识库建设。模型与训练数据已开源于https://github.com/knowlab/IHC-LLMiner。

原文摘要 · Abstract (English)

Immunohistochemistry (IHC) is essential in diagnostic pathology and biomedical research, offering critical insights into protein expression and tumour biology. This study presents an automated pipeline, IHC-LLMiner, for extracting IHC-tumour profiles from PubMed abstracts, leveraging advanced biomedical text mining. There are two subtasks: abstract classification (include/exclude as relevant) and IHC-tumour profile extraction on relevant included abstracts. The best-performing model, "Gemma-2 finetuned", achieved 91.5% accuracy and an F1 score of 91.4, outperforming GPT4-O by 9.5% accuracy with 5.9 times faster inference time. From an initial dataset of 107,759 abstracts identified for 50 immunohistochemical markers, the classification task identified 30,481 relevant abstracts (Include) using the Gemma-2 finetuned model. For IHC-tumour profile extraction, the Gemma-2 finetuned model achieved the best performance with 63.3% Correct outputs. Extracted IHC-tumour profiles (tumour types and markers) were normalised to Unified Medical Language System (UMLS) concepts to ensure consistency and facilitate IHC-tumour profile landscape analysis. The extracted IHC-tumour profiles demonstrated excellent concordance with available online summary data and provided considerable added value in terms of both missing IHC-tumour profiles and quantitative assessments. Our proposed LLM based pipeline provides a practical solution for large-scale IHC-tumour profile data mining, enhancing the accessibility and utility of such data for research and clinical applications as well as enabling the generation of quantitative and structured data to support cancer-specific knowledge base development. Models and training datasets are available at https://github.com/knowlab/IHC-LLMiner.

免疫组化文本挖掘大模型癌症研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。