arXiv:2508.15149cs.LG2025-08被引 1

用微调的RoBERTa模型自动提取病理报告中的癌种信息,准确率达80.61%。

A Robust BERT-Based Deep Learning Model for Automated Cancer Type Extraction from Unstructured Pathology Reports

  • 基于RoBERTa模型微调,专用于病理报告中癌种识别
  • F1_Bertscore达0.98,精确匹配率80.61%,优于基线和Mistral 7B
  • 适合精准肿瘤学研究与分子肿瘤委员会流程集成

从电子病历中准确提取临床信息对临床研究至关重要,但依赖大量专业人力。本研究开发了一种基于微调RoBERTa模型的自动化系统,用于从无结构病理报告中提取特定癌种信息,以支持精准肿瘤学研究。该模型显著优于基线模型和大型语言模型Mistral 7B,F1_Bertscore达到0.98,整体精确匹配率为80.61%。该微调方法展现出良好的可扩展性,可无缝集成至分子肿瘤委员会工作流程中。针对肿瘤学领域任务微调专用模型,有望实现更高效、准确的临床信息提取。

原文摘要 · Abstract (English)

The accurate extraction of clinical information from electronic medical records is particularly critical to clinical research but require much trained expertise and manual labor. In this study we developed a robust system for automated extraction of the specific cancer types for the purpose of supporting precision oncology research. from pathology reports using a fine-tuned RoBERTa model. This model significantly outperformed the baseline model and a Large Language Model, Mistral 7B, achieving F1_Bertscore 0.98 and overall exact match of 80.61%. This fine-tuning approach demonstrates the potential for scalability that can integrate seamlessly into the molecular tumour board process. Fine-tuning domain-specific models for precision tasks in oncology, may pave the way for more efficient and accurate clinical information extraction.

癌症类型提取RoBERTa病理报告精准医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。