arXiv:2603.17041stat.MLcs.AI2026-03中稿 · MathAI 2026被引 1

提出用协方差一致性评估生成模型的联合结构保真度,弥补仅看单变量分布的不足。

When Marginals Match but Structure Fails: Covariance Fidelity in Generative Models

  • 引入协方差距离衡量生成数据与真实数据的联合依赖结构差异
  • 发现即使单变量分布完全匹配,协方差差异仍可极大,导致下游分析失稳
  • 适用于对依赖结构敏感的任务,如主成分分析,且在多领域验证有效

生成模型日益用于替代真实数据开展下游科学分析,但现有评估标准仍聚焦于单变量分布匹配。我们指出这一做法存在根本性缺陷:下游推断极少是单变量操作,一个通过所有单变量检验的模型仍可能产生结构不可靠的合成数据。为此提出协方差级依赖保真度,以 D_Sigma(P,Q) = ||Sigma_P - Sigma_Q||_F 为可计算准则,评估生成模型是否保留了数据的联合结构。三个理论结果证实该准则:第一,边际保真对依赖结构无约束,D_Sigma 可无限大而单变量分布完全匹配;第二,协方差发散会引发可量化的下游不稳定性,包括总体回归系数符号反转;第三,控制 D_Sigma 可对依赖敏感任务(如 PCA)提供类似 Davis-Kahan 的正向稳定性保证。在三个领域实证验证:图像数据(Fashion-MNIST VAE, n=60,000)、bulk RNA-seq(TCGA-BRCA, n=1,111)及小样本压力测试(阿尔茨海默病基因表达, n=113),结果显示 D_Sigma/delta 在标准边际诊断难以区分时,仍能有效识别结构丢失与结构保持的生成器,证明其信息独立于传统指标,跨领域和样本量均有效。

原文摘要 · Abstract (English)

Generative models are increasingly deployed as substitutes for real data in downstream scientific workflows, yet standard evaluation criteria remain focused on marginal distribution matching. We argue that this represents a fundamental gap: downstream inference is rarely a marginal operation, and a model that passes every univariate diagnostic can still produce structurally unreliable synthetic data. We introduce covariance-level dependence fidelity, measured by D_Sigma(P,Q) = ||Sigma_P - Sigma_Q||_F, as a principled, computable criterion for evaluating whether a generative model preserves the joint structure of data beyond its univariate marginals. Three results formalise this criterion. First, marginal fidelity provides no constraint on dependence structure: D_Sigma can be made arbitrarily large while all univariate marginals match exactly. Second, covariance divergence induces quantifiable downstream instability, including sign reversals in population regression coefficients. Third, bounding D_Sigma provides positive stability guarantees for dependence-sensitive procedures such as PCA via Davis-Kahan-type bounds. Empirical validation across three domains, image data (Fashion-MNIST VAE, n = 60,000), bulk RNA-seq (TCGA-BRCA, n = 1,111), and a small-sample stress test (Alzheimer's gene expression, n = 113), shows that D_Sigma/delta consistently distinguishes structure-discarding from structure-preserving generators in cases where standard marginal diagnostics show little separation, confirming that covariance-level fidelity provides information orthogonal to existing evaluation metrics across domains and sample sizes.

生成模型协方差数据质量评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。