用患者级标签训练肿瘤分类模型,无需逐份标注报告。
Learning from Lost Provenance: Multiple Instance Learning for Cancer Registry Tumor Group Classification

- 用注意力多实例学习恢复患者标签与病理报告的关联
- 精炼后数据集使分类器宏F1达0.83,优于多数基线
- 适合想低成本自动化癌症登记工作的研究者
将深度学习应用于癌症登记系统,可自动完成病理科报告编码等繁琐任务。然而,受限于报告级人工标注数据稀缺。癌症登记机构日常生成大量专家分配的患者级标签,但这些标签未与具体病历报告绑定,难以直接用于模型训练。本文提出一种高效框架,利用操作中产生的患者级标签进行深度学习模型训练,无需逐份人工标注。以不列颠哥伦比亚癌症登记处为例,采用基于注意力的多实例学习(ABMIL)方法,通过模型对各报告的关注度,从大量噪声标签中提炼出高质量的报告级训练数据。在精炼数据集上微调的分类器达到0.83的宏F1,多数肿瘤组表现优于现有基线。该方法将常规运营标签转化为高质训练数据,无需额外标注或大规模计算资源,为自动化癌症登记流程提供可行路径。
原文摘要 · Abstract (English)
Modernizing cancer registries with deep learning is opening new opportunities to automate labor-intensive tasks such as the coding of pathology reports. However, progress is constrained by the scarcity of report-level human-annotated training data. Cancer registries generate substantial volumes of expert-assigned labels as a routine product of their operations, but these exist at the patient level and are not linked to the individual pathology reports that informed them, limiting their direct use for training models. We develop an efficient framework for training deep learning classifiers by leveraging these operationally-generated labels without requiring per-report human annotation, demonstrated for tumor group classification at the BC Cancer Registry. We use Attention-Based Multiple Instance Learning (ABMIL) to recover the lost link between patient-level labels and the reports that informed them, leveraging the attention the model places on each report to distil a large, noisily-labeled corpus into a compact, high-quality per-report training dataset. A classifier fine-tuned on a distilled dataset achieved a macro F1 of 0.83, outperforming established baselines across most tumor groups. By turning routine operational labels into high-quality training data without additional annotation or large-scale computing infrastructure, ABMIL offers a practical and accessible route to automating cancer registry workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。