构建多语言匿名化基准,支持跨语言医疗数据隐私保护研究
MultiGraSCCo: A Multilingual Anonymization Benchmark with Annotations of Personal Identifiers
- 用机器翻译生成十种语言的匿名化数据,保留原始标注并适配文化语境
- 涵盖2500+个人标识符标注,经医生验证翻译质量高
- 适合用于训练标注员、跨机构数据验证及自动识别模型优化
由于隐私顾虑,获取敏感患者数据用于机器学习存在挑战。包含个人身份信息标注的数据集对于开发和测试匿名化系统至关重要,可确保符合隐私法规的安全数据共享。由于真实患者数据难以获取,合成数据成为缓解数据稀缺的有效方案,且绕过真实数据相关的隐私监管。此外,神经机器翻译可通过将高资源语言中已验证的真实或合成数据翻译到低资源语言,生成高质量数据。本文构建了一个涵盖十种语言的多语言匿名化基准,采用保持原始标注并使地名、人名在目标语言中文化语境恰当的翻译方法。对医疗专业人士的评估证实了翻译质量,包括一般性和个人信息翻译适应性。该基准包含超过2500个个人标识符标注,可用于训练标注员、无法律障碍的跨机构标注验证,以及提升自动个人身份信息检测性能。我们公开了基准数据与标注指南,以促进后续研究。
原文摘要 · Abstract (English)
Accessing sensitive patient data for machine learning is challenging due to privacy concerns. Datasets with annotations of personally identifiable information are crucial for developing and testing anonymization systems to enable safe data sharing that complies with privacy regulations. Since accessing real patient data is a bottleneck, synthetic data offers an efficient solution for data scarcity, bypassing privacy regulations that apply to real data. Moreover, neural machine translation can help to create high-quality data for low-resource languages by translating validated real or synthetic data from a high-resource language. In this work, we create a multilingual anonymization benchmark in ten languages, using a machine translation methodology that preserves the original annotations and renders names of cities and people in a culturally and contextually appropriate form in each target language. Our evaluation study with medical professionals confirms the quality of the translations, both in general and with respect to the translation and adaptation of personal information. Our benchmark with over 2,500 annotations of personal information can be used in many applications, including training annotators, validating annotations across institutions without legal complications, and helping improve the performance of automatic personal information detection. We make our benchmark and annotation guidelines available for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。