用大模型自动脱敏医疗数据,兼顾隐私与可用性。
RedactOR: An LLM-Powered Framework for Automatic Clinical Data De-Identification
- 结合规则与大模型的混合方法,智能路由处理不同类型数据
- 在i2b2数据集上召回率达标,且降低大模型调用成本
- 适合医疗AI落地中的数据安全需求,尤其音频记录脱敏
保障临床数据隐私的同时保持数据价值,对驱动医疗AI和数据分析至关重要。现有去标识化(De-ID)方法包括基于规则的技术、深度学习模型和大语言模型(LLMs),常面临召回错误、泛化能力差和效率低等问题,限制了实际应用。我们提出一个完全自动化的多模态框架RedactOR,用于结构化与非结构化电子健康记录(含临床语音记录)的去标识化。该框架采用低成本的去标识策略,包括智能路由、规则与大模型结合的方法,以及两步式语音脱敏流程。提出基于检索的实体重词汇化方法,确保受保护实体替换的一致性,提升下游应用的数据连贯性。讨论了红岩架构的关键设计原则、去标识化与重词汇化方法及模块化结构,并集成至Oracle Health临床AI系统。在i2b2 2014去标识化数据集上,以标准指标评估,本方法表现优异,同时优化了令牌使用,降低大模型成本。最后分享了在真实医疗AI数据管道中部署的经验与洞察。
原文摘要 · Abstract (English)
Ensuring clinical data privacy while preserving utility is critical for AI-driven healthcare and data analytics. Existing de-identification (De-ID) methods, including rule-based techniques, deep learning models, and large language models (LLMs), often suffer from recall errors, limited generalization, and inefficiencies, limiting their real-world applicability. We propose a fully automated, multi-modal framework, RedactOR for de-identifying structured and unstructured electronic health records, including clinical audio records. Our framework employs cost-efficient De-ID strategies, including intelligent routing, hybrid rule and LLM based approaches, and a two-step audio redaction approach. We present a retrieval-based entity relexicalization approach to ensure consistent substitutions of protected entities, thereby enhancing data coherence for downstream applications. We discuss key design desiderata, de-identification and relexicalization methodology, and modular architecture of RedactOR and its integration with the Oracle Health Clinical AI system. Evaluated on the i2b2 2014 De-ID dataset using standard metrics with strict recall, our approach achieves competitive performance while optimizing token usage to reduce LLM costs. Finally, we discuss key lessons and insights from deployment in real-world AI- driven healthcare data pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。