arXiv:2601.00787cs.CL2026-01

用少量微调让癌症病理模型跨省通用,显著减少漏诊。

Adapting Natural Language Processing Models Across Jurisdictions: A pilot Study in Canadian Cancer Registries

  • 用加拿大不同省份的病理报告微调模型,实现跨区域迁移。
  • 集成双模型后漏诊数减半,关键任务召回率达99%。
  • 仅共享模型权重,保护隐私,支持全国统一病历分析系统。

基于人群的癌症登记依赖病理报告作为主要诊断来源,但人工提取数据耗时费力,导致数据延迟。尽管基于Transformer的NLP系统已提升登记效率,其在不同地区间因报告格式差异导致的泛化能力仍不明确。本研究首次评估了在加拿大跨省使用由卑诗省癌症登记处开发的领域自适应模型BCCRTron,以及生物医学Transformer模型GatorTron进行癌症监测的表现。训练数据来自纽芬兰与拉布拉多癌症登记处(NLCR),包含约10.4万份用于一级分类(癌症/非癌症)和2.2万份用于二级分类(可报告/不可报告)的脱敏病理报告。两种模型分别通过结构化报告段落与诊断文本输入管道进行微调。在NLCR测试集上,经适配的模型保持高性能,证明预训练于一地区的Transformer可通过小规模微调迁移到另一地区。为提高敏感性,采用保守或组合策略集成两模型,一级任务召回率达0.99,漏诊数降至24例(原模型分别为48和54);二级任务召回率同样达0.99,漏报可报告病例数降至33例(原模型分别为54和46)。结果表明,结合互补文本表示的集成模型能显著减少漏诊,提升错误覆盖。研究还提出一种隐私保护流程,仅共享模型权重,促进各省份间NLP系统的互操作性,并为未来建立全国性癌症病理与登记通用模型奠定基础。

原文摘要 · Abstract (English)

Population-based cancer registries depend on pathology reports as their primary diagnostic source, yet manual abstraction is resource-intensive and contributes to delays in cancer data. While transformer-based NLP systems have improved registry workflows, their ability to generalize across jurisdictions with differing reporting conventions remains poorly understood. We present the first cross-provincial evaluation of adapting BCCRTron, a domain-adapted transformer model developed at the British Columbia Cancer Registry, alongside GatorTron, a biomedical transformer model, for cancer surveillance in Canada. Our training dataset consisted of approximately 104,000 and 22,000 de-identified pathology reports from the Newfoundland & Labrador Cancer Registry (NLCR) for Tier 1 (cancer vs. non-cancer) and Tier 2 (reportable vs. non-reportable) tasks, respectively. Both models were fine-tuned using complementary synoptic and diagnosis focused report section input pipelines. Across NLCR test sets, the adapted models maintained high performance, demonstrating transformers pretrained in one jurisdiction can be localized to another with modest fine-tuning. To improve sensitivity, we combined the two models using a conservative OR-ensemble achieving a Tier 1 recall of 0.99 and reduced missed cancers to 24, compared with 48 and 54 for the standalone models. For Tier 2, the ensemble achieved 0.99 recall and reduced missed reportable cancers to 33, compared with 54 and 46 for the individual models. These findings demonstrate that an ensemble combining complementary text representations substantially reduce missed cancers and improve error coverage in cancer-registry NLP. We implement a privacy-preserving workflow in which only model weights are shared between provinces, supporting interoperable NLP infrastructure and a future pan-Canadian foundation model for cancer pathology and registry workflows.

癌症登记NLP迁移隐私保护模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。