arXiv:2504.18596cs.CRcs.AI2025-04被引 2

用合成数据和可配置扰动提升隐私与数据效用的平衡,适合金融等敏感行业

Optimizing the Privacy-Utility Balance using Synthetic Data and Configurable Perturbation Pipelines

  • 结合生成模型与可配置扰动,构建高保真隐私数据集
  • 相比传统匿名化,显著提升数据可用性与安全性
  • 适合金融、医疗等强监管领域做安全数据分析

本文探讨现代合成数据生成与先进数据扰动技术在管理大规模数据集中的战略应用,重点面向银行、金融与保险(BFSI)行业。对比生成对抗网络(GANs)、上下文感知的个人身份信息(PII)转换、可配置统计扰动及差分隐私等方法与传统匿名化手段,旨在创建既真实又具备隐私保护能力的数据集,以支持复杂机器学习任务与分析需求。这些技术有望在保障隐私的同时大幅提高数据效用,并带来运营优势,如降低处理开销、加快分析速度。研究还揭示其在降低合规风险、推动可扩展数据驱动创新方面的潜力,实现敏感客户信息不泄露前提下的业务发展。

原文摘要 · Abstract (English)

This paper explores the strategic use of modern synthetic data generation and advanced data perturbation techniques to enhance security, maintain analytical utility, and improve operational efficiency when managing large datasets, with a particular focus on the Banking, Financial Services, and Insurance (BFSI) sector. We contrast these advanced methods encompassing generative models like GANs, sophisticated context-aware PII transformation, configurable statistical perturbation, and differential privacy with traditional anonymization approaches. The goal is to create realistic, privacy-preserving datasets that retain high utility for complex machine learning tasks and analytics, a critical need in the data-sensitive industries like BFSI, Healthcare, Retail, and Telecommunications. We discuss how these modern techniques potentially offer significant improvements in balancing privacy preservation while maintaining data utility compared to older methods. Furthermore, we examine the potential for operational gains, such as reduced overhead and accelerated analytics, by using these privacy-enhanced datasets. We also explore key use cases where these methods can mitigate regulatory risks and enable scalable, data-driven innovation without compromising sensitive customer information.

合成数据隐私保护数据扰动BFSI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。