用概率关系模型生成隐私安全的合成关系数据
Towards Privacy-Preserving Relational Data Synthesis via Probabilistic Relational Models
- 从关系数据库构建概率关系模型,实现数据生成
- 可基于模型采样新合成数据点,保持原始关系结构
- 适合需要隐私保护数据的机器学习研究者
概率关系模型将一阶逻辑与概率模型结合,能有效表示关系域中对象之间的关联。然而,人工智能任务日益依赖大规模关系型训练数据,而真实数据收集常受隐私担忧、法规限制和高成本制约。为此,生成合成数据成为可行方案。本文提出一个完整的端到端流程,将关系数据库转化为概率关系模型,并基于其概率分布采样生成新的合成关系数据。核心贡献包括:设计一种学习算法,从给定关系数据库构建概率关系模型,从而实现对关系结构的精准建模与隐私保护下的数据合成。
原文摘要 · Abstract (English)
Probabilistic relational models provide a well-established formalism to combine first-order logic and probabilistic models, thereby allowing to represent relationships between objects in a relational domain. At the same time, the field of artificial intelligence requires increasingly large amounts of relational training data for various machine learning tasks. Collecting real-world data, however, is often challenging due to privacy concerns, data protection regulations, high costs, and so on. To mitigate these challenges, the generation of synthetic data is a promising approach. In this paper, we solve the problem of generating synthetic relational data via probabilistic relational models. In particular, we propose a fully-fledged pipeline to go from relational database to probabilistic relational model, which can then be used to sample new synthetic relational data points from its underlying probability distribution. As part of our proposed pipeline, we introduce a learning algorithm to construct a probabilistic relational model from a given relational database.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。