arXiv:2504.20657cs.CV2025-04

基于XNAT平台构建医疗影像去标识化流程,提升数据隐私保护能力。

Image deidentification in the XNAT ecosystem: use cases and solutions

  • 利用XNAT内置工具与生态独立组件,设计可复用的DICOM去标识化工作流。
  • 规则方法实现99.61%去标识准确率,但对地址信息处理仍有不足。
  • 适用于医学影像研究团队,尤其关注数据共享与合规性的项目。

XNAT是学术界广泛使用的基于服务器的数据管理平台,用于整理大型DICOM影像数据库。本文详述了在XNAT生态系统中使用平台功能及配套工具实现DICOM数据去标识化的流程。结合过往经验,列出多种需去标识的使用场景。参与医学影像去标识基准挑战(MIDI-B)的初始方法源于本地已有方案,经验证阶段调整后,在测试阶段取得97.91%准确率,因与赛方Synapse平台存在技术不兼容问题,导致无法获取反馈。提交后,通过组织方提供的差异报告和连续基准测试机制,性能提升至99.61%。纯规则方法可完全移除姓名相关数据,但在地址识别上存在缺陷。使用公开机器学习模型处理地址部分部分成功,但对其他自由文本过度删除,造成整体性能下降至99.54%。未来将聚焦于提升地址识别能力,并改进图像像素中嵌入的可识别信息去除。目前估计真实去标识失败率为0.19%,关于“答案键”的技术细节仍在与组织方讨论中。

原文摘要 · Abstract (English)

XNAT is a server-based data management platform widely used in academia for curating large databases of DICOM images for research projects. We describe in detail a deidentification workflow for DICOM data using facilities in XNAT, together with independent tools in the XNAT "ecosystem". We list different contexts in which deidentification might be needed, based on our prior experience. The starting point for participation in the Medical Image De-Identification Benchmark (MIDI-B) challenge was a set of pre-existing local methodologies, which were adapted during the validation phase of the challenge. Our result in the test phase was 97.91\%, considerably lower than our peers, due largely to an arcane technical incompatibility of our methodology with the challenge's Synapse platform, which prevented us receiving feedback during the validation phase. Post-submission, additional discrepancy reports from the organisers and via the MIDI-B Continuous Benchmarking facility, enabled us to improve this score significantly to 99.61\%. An entirely rule-based approach was shown to be capable of removing all name-related information in the test corpus, but exhibited failures in dealing fully with address data. Initial experiments using published machine-learning models to remove addresses were partially successful but showed the models to be "over-aggressive" on other types of free-text data, leading to a slight overall degradation in performance to 99.54\%. Future development will therefore focus on improving address-recognition capabilities, but also on better removal of identifiable data burned into the image pixels. Several technical aspects relating to the "answer key" are still under discussion with the challenge organisers, but we estimate that our percentage of genuine deidentification failures on the MIDI-B test corpus currently stands at 0.19\%. (Abridged from original for arXiv submission)

医疗影像去标识化数据安全XNAT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。