用逻辑求解生成私密且真实的合成基因数据,效率远超现有方法。
Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers
- 基于SAT求解器构建生成框架,确保数据逻辑一致与隐私保护。
- 相比主流方法,准确率提升84%-93%,隐私保护达95%-98%提升。
- 可调节隐私与精度平衡,适合从敏感群体到临床应用的多种场景。
机器生成的数据对训练人工智能、评估罕见工作流及在严格数据法规下共享数据具有重要价值。当前统计与深度学习方法在大规模数据上表现不佳,易生成不真实场景,且难以量化隐私。本文提出Genomator,一种基于逻辑求解(SAT求解)的方法,高效生成私密且符合现实的数据表示。我们在基因组数据上验证该方法,这类数据是复杂度与隐私性最高的信息之一。合成基因组有望平衡医学研究中代表性不足的人群,并推动全球数据共享。在基准测试中,Genomator相较于马尔可夫生成、受限玻尔兹曼机、生成对抗网络及条件受限玻尔兹曼机,准确率提升84%-93%,隐私保护能力提高95%-98%。其效率高出1000-1600倍,是唯一可扩展至全基因组的测试方法。我们揭示了隐私与准确率间的普遍权衡,并利用Genomator的可调性适配不同应用场景,从可证明私密的敏感人群表示,到难以区分的药理基因组特征数据集。展示可扩展的可调合成数据生成,有助于增强信任并推动其进入临床应用。
原文摘要 · Abstract (English)
Machine-generated data is a valuable resource for training Artificial Intelligence algorithms, evaluating rare workflows, and sharing data under stricter data legislations. The challenge is to generate data that is accurate and private. Current statistical and deep learning methods struggle with large data volumes, are prone to hallucinating scenarios incompatible with reality, and seldom quantify privacy meaningfully. Here we introduce Genomator, a logic solving approach (SAT solving), which efficiently produces private and realistic representations of the original data. We demonstrate the method on genomic data, which arguably is the most complex and private information. Synthetic genomes hold great potential for balancing underrepresented populations in medical research and advancing global data exchange. We benchmark Genomator against state-of-the-art methodologies (Markov generation, Restricted Boltzmann Machine, Generative Adversarial Network and Conditional Restricted Boltzmann Machines), demonstrating an 84-93% accuracy improvement and 95-98% higher privacy. Genomator is also 1000-1600 times more efficient, making it the only tested method that scales to whole genomes. We show the universal trade-off between privacy and accuracy, and use Genomator's tuning capability to cater to all applications along the spectrum, from provable private representations of sensitive cohorts, to datasets with indistinguishable pharmacogenomic profiles. Demonstrating the production-scale generation of tuneable synthetic data can increase trust and pave the way into the clinic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。