arXiv:2510.01793cs.LG2025-10

检验合成数据隐私过滤器,发现其无法可靠识别生成的相似图像。

Sensitivity, Specificity, and Consistency: A Tripartite Evaluation of Privacy Filters for Synthetic Data Generation

  • 用三重评估框架测试胸部X光合成中的隐私过滤器
  • 过滤器对真实图像敏感但对生成近似图像检测率低
  • 适合关注医疗数据隐私安全的研究者参考

生成保护隐私的合成数据是解决医学AI研究中数据稀缺问题的有前景方向。近期提出的后处理隐私过滤技术旨在移除包含个人身份信息的样本,但其有效性尚未得到充分验证。本文针对胸部X光合成场景,对一个过滤流程进行了严格评估。与原始论文声称相反,我们的结果表明当前过滤器在特异性和一致性方面表现有限,仅对真实图像具有高敏感性,却无法可靠检测由训练数据生成的近似重复图像。这揭示了后处理过滤的一个关键缺陷:不仅未能有效保障患者隐私,反而可能提供虚假的安全感,导致不可接受的信息泄露。我们结论是,要使此类方法在敏感应用中可信部署,需在过滤器设计上实现重大改进。

原文摘要 · Abstract (English)

The generation of privacy-preserving synthetic datasets is a promising avenue for overcoming data scarcity in medical AI research. Post-hoc privacy filtering techniques, designed to remove samples containing personally identifiable information, have recently been proposed as a solution. However, their effectiveness remains largely unverified. This work presents a rigorous evaluation of a filtering pipeline applied to chest X-ray synthesis. Contrary to claims from the original publications, our results demonstrate that current filters exhibit limited specificity and consistency, achieving high sensitivity only for real images while failing to reliably detect near-duplicates generated from training data. These results demonstrate a critical limitation of post-hoc filtering: rather than effectively safeguarding patient privacy, these methods may provide a false sense of security while leaving unacceptable levels of patient information exposed. We conclude that substantial advances in filter design are needed before these methods can be confidently deployed in sensitive applications.

隐私保护合成数据医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。