通过粗化边缘数据生成可信赖的合成数据,确保关系保留且无泄露风险。
It does what it says on the tin: safe synthetic data from coarsened margins

- 用粗化边缘数据构建合成数据,明确保留原始数据中的变量关系。
- 经统计披露控制处理的边缘数据用于生成合成数据,确保无泄露风险。
- 适合需要高透明度与安全性的政府或机构数据共享场景。
本文提出一种生成合成数据(SD)的方法,具有两大优势:一是透明性,接收者可明确知晓原始数据中哪些变量关系将在合成数据中近似保留;二是安全性,合成数据仅基于已被认定无披露风险的信息生成。方法首先定义并计算将保留关系的边缘数据,随后对各边缘实施统计披露控制(SDC),如顶编码、底编码、小类别合并及小计数调整等,符合数据保管方标准。进一步通过将表格中所有计数粗化为披露阈值的倍数进行优化调整。最终使用迭代比例拟合(IPF)算法基于这些处理后的边缘数据生成合成数据。文中以1901年苏格兰人口普查数据为例,演示了实际操作步骤。
原文摘要 · Abstract (English)
This paper proposes a method of creating synthetic data (SD) that will have two important advantages for the user compared to other methods currently available. The first is transparency; unlike other methods, the person in receipt of the SD will know which of the relationships between variables in the original data will be approximately maintained in the SD. The second is a guarantee that the SD is derived from information that has already been judged to be free of disclosure risk. This is achieved by first defining and calculating the margins where relationships between variables will be maintained in the SD. Each margin will then be subject to statistical disclosure control (SDC) to the standards defined by the data custodian, e.g. top-coding and bottom-coding, combination of small categories and/or modifying small counts. Further adjustment of the curated margins is advised by coarsening all counts in the table to multiples of the disclosure limit. These adjusted margins are used to create SD by the Iterative Proportional Fitting (IPF) algorithm. The practical steps involved in creating such SD are illustrated using data from the 1901 Census of Scotland.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。