arXiv:2509.14464cs.CL2025-09EMNLP被引 8

评估大模型在医疗去标识化中的信息丢失问题,发现现有方法存在严重误判。

Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss

  • 系统梳理大模型医疗去标识研究,揭示报告标准不统一
  • 实测多模型发现临床信息被错误删除率超30%
  • 引入专家人工验证,提出新方法识别关键信息误删

医疗去标识化是自然语言处理在医疗领域的应用,旨在自动移除患者(有时也包括医护人员)的个人身份信息。随着生成式大语言模型(LLMs)的兴起,大量研究尝试将其应用于去标识化任务。尽管这些方法常报告接近完美的性能,但其可复现性和实用性仍面临严峻挑战。本文识别出当前文献中的三大局限:报告指标不一致导致难以直接比较;传统分类指标无法有效捕捉大模型易犯的错误(如篡改临床相关信息);缺乏对自动化评估指标的人工验证,而这些指标旨在量化此类错误。为此,我们首先开展一项关于基于大模型的去标识化研究的调查,揭示报告标准的异质性。其次,评估多种模型以量化临床信息被不当移除的程度。接着,通过临床专家进行人工验证,评估现有评估指标在检测临床信息移除方面的有效性,揭示其性能不佳及内在局限性。最后,我们提出一种新型方法,用于检测临床相关信息的误删。

原文摘要 · Abstract (English)

De-identification in the healthcare setting is an application of NLP where automated algorithms are used to remove personally identifying information of patients (and, sometimes, providers). With the recent rise of generative large language models (LLMs), there has been a corresponding rise in the number of papers that apply LLMs to de-identification. Although these approaches often report near-perfect results, significant challenges concerning reproducibility and utility of the research papers persist. This paper identifies three key limitations in the current literature: inconsistent reporting metrics hindering direct comparisons, the inadequacy of traditional classification metrics in capturing errors which LLMs may be more prone to (i.e., altering clinically relevant information), and lack of manual validation of automated metrics which aim to quantify these errors. To address these issues, we first present a survey of LLM-based de-identification research, highlighting the heterogeneity in reporting standards. Second, we evaluated a diverse set of models to quantify the extent of inappropriate removal of clinical information. Next, we conduct a manual validation of an existing evaluation metric to measure the removal of clinical information, employing clinical experts to assess their efficacy. We highlight poor performance and describe the inherent limitations of such metrics in identifying clinically significant changes. Lastly, we propose a novel methodology for the detection of clinically relevant information removal.

医疗AI大模型去标识化信息丢失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。