arXiv:2510.21610cs.LGcs.AI2025-10

用数学方法生成保留复杂关联的合成数据,解决隐私与模型训练难题

Generative Correlation Manifolds: Generating Synthetic Data with Preserved Higher-Order Correlations

  • 基于目标相关矩阵的Cholesky分解,生成保持完整相关结构的数据
  • 理论上可精确复现从两两关系到高阶交互的所有相关性
  • 适合需要真实数据复杂结构的隐私保护、模型训练场景

数据隐私需求和对鲁棒机器学习模型的追求推动了合成数据生成技术的发展。然而,现有方法通常只能复制简单的统计特征,难以保留真实世界系统中复杂的多变量交互所依赖的成对及高阶相关结构。这一局限可能导致合成数据看似真实,但在复杂建模任务中失效。本文提出生成相关流形(Generative Correlation Manifolds, GCM),一种计算高效的合成数据生成方法。该方法利用目标相关矩阵的Cholesky分解,生成的数据在数学上被证明能完全保留源数据的全部相关结构——从简单的成对关系到高阶交互。我们主张该方法为合成数据生成提供了新范式,具有在隐私保护数据共享、鲁棒模型训练和仿真中的潜在应用价值。

原文摘要 · Abstract (English)

The increasing need for data privacy and the demand for robust machine learning models have fueled the development of synthetic data generation techniques. However, current methods often succeed in replicating simple summary statistics but fail to preserve both the pairwise and higher-order correlation structure of the data that define the complex, multi-variable interactions inherent in real-world systems. This limitation can lead to synthetic data that is superficially realistic but fails when used for sophisticated modeling tasks. In this white paper, we introduce Generative Correlation Manifolds (GCM), a computationally efficient method for generating synthetic data. The technique uses Cholesky decomposition of a target correlation matrix to produce datasets that, by mathematical proof, preserve the entire correlation structure -- from simple pairwise relationships to higher-order interactions -- of the source dataset. We argue that this method provides a new approach to synthetic data generation with potential applications in privacy-preserving data sharing, robust model training, and simulation.

合成数据相关结构隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。