用差分隐私生成数据,既保护隐私又保留金融数据价值。
Decoupling Identity from Utility: Privacy-by-Design Frameworks for Financial Ecosystems
- 用差分隐私合成数据,分离身份与用途,避免重识别风险。
- 表格合成适合静态数据分析,代理模型可模拟市场黑天鹅事件。
- 适合金融机构合规研究和跨机构数据协作,符合严格监管要求。
金融机构在提升数据效用与降低重识别风险之间面临矛盾。本文探索差分隐私(DP)合成数据作为“隐私设计”框架,可在保障输出隐私的同时满足严格监管要求。研究对比两种生成范式:直接表格合成,从原始数据重建高保真联合分布,适用于问答测试与业务分析;以及基于差分隐私种子的代理模型(DP-Seeded ABM),利用受保护的聚合数据参数化复杂动态仿真。前者反映历史静态相关性,后者提供可模拟市场动态行为与极端事件的“反事实实验室”。通过将个体身份与数据效用解耦,该方法消除传统数据清理瓶颈,支持跨机构研究与合规决策,在不断演变的监管环境中实现数据安全共享。
原文摘要 · Abstract (English)
Financial institutions face tension between maximizing data utility and mitigating the re-identification risks inherent in traditional anonymization methods. This paper explores Differentially Private (DP) synthetic data as a robust "Privacy by Design" framework to resolve this conflict, ensuring output privacy while satisfying stringent regulatory obligations. We examine two distinct generative paradigms: Direct Tabular Synthesis, which reconstructs high-fidelity joint distributions from raw data, and DP-Seeded Agent-Based Modeling (ABM), which uses DP-protected aggregates to parameterize complex, stateful simulations. While tabular synthesis excels at reflecting static historical correlations for QA testing and business analytics, the DP-Seeded ABM offers a forward-looking "counterfactual laboratory" capable of modeling dynamic market behaviors and black swan events. By decoupling individual identities from data utility, these methodologies eliminate traditional data-clearing bottlenecks, enabling seamless cross-institutional research and compliant decision-making in an evolving regulatory landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。