arXiv:2603.27254cs.DBcs.AI2026-03

融合大模型与概率图模型,生成既真实又适合分析的合成数据。

Amalgam: Hybrid LLM-PGM Synthesis Algorithm for Accuracy and Realism

  • 结合LLM的复杂结构生成能力与PGM的统计准确性。
  • 合成数据平均χ² P值达91%,真实感评分3.8/5。
  • 适合医疗等需要高真实性和隐私保护的领域使用。

为生成合成数据(如医疗领域),现有方法主要分为概率图模型(PGMs)和深度学习模型(如LLMs)。PGMs生成的数据可用于高级分析,但难以支持复杂模式;而LLMs虽能处理复杂结构,却导致数据分布失真,影响分析效果。本文提出Amalgam,一种混合式LLM-PGM数据合成算法,兼顾高级分析能力、数据真实性和可量化的隐私属性。实验显示,Amalgam生成数据的平均χ² P值达到91%,在自研真实感评估指标下得分为3.8/5,优于当前最优水平(3.3),接近真实数据表现(4.7)。

原文摘要 · Abstract (English)

To generate synthetic datasets, e.g., in domains such as healthcare, the literature proposes approaches of two main types: Probabilistic Graphical Models (PGMs) and Deep Learning models, such as LLMs. While PGMs produce synthetic data that can be used for advanced analytics, they do not support complex schemas and datasets. LLMs on the other hand, support complex schemas but produce skewed dataset distributions, which are less useful for advanced analytics. In this paper, we therefore present Amalgam, a hybrid LLM-PGM data synthesis algorithm supporting both advanced analytics, realism, and tangible privacy properties. We show that Amalgam synthesizes data with an average 91 % $χ^2 P$ value and scores 3.8/5 for realism using our proposed metric, where state-of-the-art is 3.3 and real data is 4.7.

数据合成大模型概率图模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。