针对病理报告分类模型域外性能下降问题,提出可复现的训练与评估流水线。
In-Domain Supervised Pathology Report Classification: A Reproducible Pipeline from Data Curation to Production-Matched Evaluation

- 基于机构分层采样构建域内训练集与生产匹配测试集
- 实现0.3%假阴性率与9.7%假阳性率,F1提升至0.922
- 通过盲审发现标签噪声严重,尤其集中在罕见癌种
我们提出一种域内监督流水线,以应对监督式生物医学NLP模型在跨癌症登记系统迁移时出现的分布外性能下降问题。该流水线提供从数据整理到生产匹配评估的可复现方案,包含如何构建域内训练集和生产匹配的保留集,以及选择保持低假阴性率(FNR)同时控制人工审查工作量的操作点。流程采用机构分层采样标准化数据整理,并对关联登记病例的报告进行独立处理,还引入盲法人工审计以估计阳性病例比例和标签噪声。在41.8万份报告的保留集中,肯塔基模型实现FNR 0.003、FPR 0.097,优于西雅图训练的MOSSAIC OncoID基线(FNR 0.010,FPR 0.183),F1从0.860提升至0.922。对600份报告的盲审显示,阳性率由0.500降至0.398,表明存在显著标签噪声,且错误集中于罕见原发部位。
原文摘要 · Abstract (English)
We introduce an in-domain supervised pipeline designed to counter the out-of-distribution performance drop that hampers supervised biomedical NLP models, a problem observed when models trained on pathology reports are moved across cancer registries. Our contribution is a reproducible recipe for training a supervised classifier from routinely collected cancer registry data. It describes how to build the in-domain training set and a production-matched holdout, and to choose operating points that keep the false-negative rate (FNR) very low while keeping reviewer workload manageable. The pipeline standardizes data curation with facility-stratified sampling and separate handling of reports linked to registry cases, and includes a blinded manual audit to estimate positive-case prevalence and label noise. On a 418k-report holdout set, the Kentucky model achieved FNR 0.003 and false-positive rate (FPR) 0.097, improving over the Seattle-trained MOSSAIC OncoID baseline (FNR 0.010, FPR 0.183) and raising F1 from 0.860 to 0.922. In a blinded manual review of 600 reports, estimated positive prevalence declined from 0.500 to 0.398, indicating substantial label noise with errors concentrated in rare primary sites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。