用医生诊断理由训练模型,能微调性能但不如多标注报告有效。
Can human clinical rationales improve the performance and explainability of clinical text classification models?
- 用医生提供的诊断理由作为额外训练数据
- 在资源充足时略提性能,但有限时表现不稳
- 若重解释性,理由数据或更利于识别关键特征
AI驱动的临床文本分类对可解释的群体健康信息自动检索至关重要。本研究探究人类提供的临床理由是否可作为额外监督信号,提升基于Transformer的模型在提取原发癌部位任务中的性能与可解释性。我们分析了99,125条提供原发癌部位诊断合理解释的人类理由,并与128,649份电子病理报告一同用于训练模型。研究还引入充分性(sufficiency)作为自动筛选理由质量的指标。结果表明,在高资源场景下,理由数据可小幅提升模型性能,但在资源受限时表现不稳定;使用充分性预筛选理由也导致不一致结果。更重要的是,仅增加报告数量的模型始终优于理由增强模型。因此,若以准确率为首要目标,应优先标注更多报告而非生成理由;若强调可解释性,理由数据可能帮助模型更好识别类似理由的特征。结论:相比增加报告,使用理由数据带来的性能提升较小,可解释性仅略有改善(以平均词级理由覆盖率衡量)。
原文摘要 · Abstract (English)
AI-driven clinical text classification is vital for explainable automated retrieval of population-level health information. This work investigates whether human-based clinical rationales can serve as additional supervision to improve both performance and explainability of transformer-based models that automatically encode clinical documents. We analyzed 99,125 human-based clinical rationales that provide plausible explanations for primary cancer site diagnoses, using them as additional training samples alongside 128,649 electronic pathology reports to evaluate transformer-based models for extracting primary cancer sites. We also investigated sufficiency as a way to measure rationale quality for pre-selecting rationales. Our results showed that clinical rationales as additional training data can improve model performance in high-resource scenarios but produce inconsistent behavior when resources are limited. Using sufficiency as an automatic metric to preselect rationales also leads to inconsistent results. Importantly, models trained on rationales were consistently outperformed by models trained on additional reports instead. This suggests that clinical rationales don't consistently improve model performance and are outperformed by simply using more reports. Therefore, if the goal is optimizing accuracy, annotation efforts should focus on labeling more reports rather than creating rationales. However, if explainability is the priority, training models on rationale-supplemented data may help them better identify rationale-like features. We conclude that using clinical rationales as additional training data results in smaller performance improvements and only slightly better explainability (measured as average token-level rationale coverage) compared to training on additional reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。