无需标注数据,用AI自动分析医疗转诊理由是否合理。
An Unsupervised Natural Language Processing Pipeline for Assessing Referral Appropriateness
- 用预训练医学语料的Transformer模型无监督聚类转诊文本
- 在近90万份真实转诊记录中识别原因准确率超92%
- 适合医保监管与医院管理,助力制定更科学的诊疗规范
评估诊断转诊的合理性对提升医疗效率、减少不必要的检查至关重要。然而,当转诊原因仅以自由文本形式记录时(如意大利国家医疗体系),该任务极具挑战性。为此,本文提出一个完全无监督的自然语言处理(NLP)流程,可在不依赖标注数据的情况下提取并评估转诊理由。该流程利用在意大利医学文本上预训练的Transformer嵌入,对转诊理由进行聚类,并评估其与适宜性指南的一致性。研究分析了伦巴第大区两个完整区域数据集:2019–2021年下肢静脉彩超(ECD;n=496,971,用于开发)和结肠镜检查(FEC;n=407,949,仅用于测试)。每项数据集随机抽取1,000条样本进行人工标注以评估性能。结果表明,该流程在识别转诊理由方面表现优异(ECD:精确率92.43%,召回率83.28%;FEC:精确率93.59%,召回率92.70%),在评估合理性方面也达到高准确率(ECD:精确率93.58%,召回率91.52%;FEC:精确率94.66%,召回率93.96%)。区域层面分析揭示了不合理转诊群体及不同场景间的差异,为伦巴第大区新政策出台提供了依据。结论表明,该方法构建了一个稳健、可扩展的无监督NLP工具,能有效挖掘大规模真实世界数据价值,为公共卫生部门提供可部署的AI监测手段,支持循证决策。
原文摘要 · Abstract (English)
Objective: Assessing the appropriateness of diagnostic referrals is critical for improving healthcare efficiency and reducing unnecessary procedures. However, this task becomes challenging when referral reasons are recorded only as free text rather than structured codes, like in the Italian NHS. To address this gap, we propose a fully unsupervised Natural Language Processing (NLP) pipeline capable of extracting and evaluating referral reasons without relying on labelled datasets. Methods: Our pipeline leverages Transformer-based embeddings pre-trained on Italian medical texts to cluster referral reasons and assess their alignment with appropriateness guidelines. It operates in an unsupervised setting and is designed to generalize across different examination types. We analyzed two complete regional datasets from the Lombardy Region (Italy), covering all referrals between 2019 and 2021 for venous echocolordoppler of the lower limbs (ECD;n=496,971; development) and flexible endoscope colonoscopy (FEC; n=407,949; testing only). For both, a random sample of 1,000 referrals was manually annotated to measure performance. Results: The pipeline achieved high performance in identifying referral reasons (Prec=92.43% (ECD), 93.59% (FEC); Rec=83.28% (ECD), 92.70% (FEC)) and appropriateness (Prec=93.58% (ECD), 94.66% (FEC); Rec=91.52% (ECD), 93.96% (FEC)). At the regional level, the analysis identified relevant inappropriate referral groups and variation across contexts, findings that informed a new Lombardy Region resolution to reinforce guideline adherence. Conclusions: This study presents a robust, scalable, unsupervised NLP pipeline for assessing referral appropriateness in large, real-world datasets. It demonstrates how such data can be effectively leveraged, providing public health authorities with a deployable AI tool to monitor practices and support evidence-based policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。