小模型在临床去标识化上效率高且泛化能力强,适合多语言场景。
Towards Fair and Efficient De-identification: Quantifying the Efficiency and Generalizability of De-identification Approaches
- 用小模型微调,在多语言、跨文化数据上表现优于大模型。
- 小模型推理成本低,性能与大模型相当,部署更实际。
- 发布BERT-MultiCulture-DEID模型,支持中、印、西、法等多语种去标识化。
大型语言模型在临床去标识化任务中表现优异,但以往研究未评估其在不同格式、文化及性别背景下的泛化能力。本文系统评估了微调的Transformer模型(BERT、ClinicalBERT、ModernBERT)、小型LLM(Llama 1-8B、Qwen 1.5-7B)和大型LLM(Llama-70B、Qwen-72B)在该任务中的表现。结果表明,小模型在保持高精度的同时显著降低推理开销,更具实用性。此外,仅用有限数据微调的小模型在识别中文、印地语、西班牙语、法语、孟加拉语及英语地区变体,以及性别化姓名方面,表现超越大模型。为提升多文化环境下的鲁棒性,本文提出并公开发布BERT-MultiCulture-DEID,基于BERT、ClinicalBERT和ModernBERT,在MIMIC数据集上对多种语言变体的标识符进行微调。本研究首次量化了去标识化中效率与泛化性的权衡关系,为公平高效的临床数据脱敏提供了可行路径。模型获取详情见:https://doi.org/10.5281/zenodo.18342291
原文摘要 · Abstract (English)
Large language models (LLMs) have shown strong performance on clinical de-identification, the task of identifying sensitive identifiers to protect privacy. However, previous work has not examined their generalizability between formats, cultures, and genders. In this work, we systematically evaluate fine-tuned transformer models (BERT, ClinicalBERT, ModernBERT), small LLMs (Llama 1-8B, Qwen 1.5-7B), and large LLMs (Llama-70B, Qwen-72B) at de-identification. We show that smaller models achieve comparable performance while substantially reducing inference cost, making them more practical for deployment. Moreover, we demonstrate that smaller models can be fine-tuned with limited data to outperform larger models in de-identifying identifiers drawn from Mandarin, Hindi, Spanish, French, Bengali, and regional variations of English, in addition to gendered names. To improve robustness in multi-cultural contexts, we introduce and publicly release BERT-MultiCulture-DEID, a set of de-identification models based on BERT, ClinicalBERT, and ModernBERT, fine-tuned on MIMIC with identifiers from multiple language variants. Our findings provide the first comprehensive quantification of the efficiency-generalizability trade-off in de-identification and establish practical pathways for fair and efficient clinical de-identification. Details on accessing the models are available at: https://doi.org/10.5281/zenodo.18342291
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。