arXiv:2412.18041stat.MLcs.LG2024-12被引 2

用信息论证明:生成数据可放大样本量而不增加信息,但有严格数学上限。

An information theoretic limit to data amplification

  • 基于信息论构建数据放大上限,仅依赖训练与生成事件数。
  • 证明增益大于1可行,但变量分辨率不提升,仅增强统计显著性。
  • 适用于需高效生成数据的科学计算场景,如高能物理模拟。

近年来,生成式人工智能被用于支持科学分析的数据生成。例如,生成对抗网络(GAN)在蒙特卡洛模拟输入上训练后,可生成同类型数据,显著降低计算时间。训练N个事件的GAN可生成GN个生成事件,增益因子G超过1。这看似违反了‘信息无法免费获取’原则。本文将此过程称为数据放大,并通过信息论概念进行研究。结果表明,在保持数据信息量不变的前提下,增益大于1是可能的,且存在仅依赖生成与训练事件数的数学边界。研究给出了原始与重构概率分布需满足的条件。特别地,放大过程不提高变量分辨率,但样本量增加仍可提升统计显著性。该边界通过计算机模拟及文献中GAN生成数据的分析得到验证。

原文摘要 · Abstract (English)

In recent years generative artificial intelligence has been used to create data to support science analysis. For example, Generative Adversarial Networks (GANs) have been trained using Monte Carlo simulated input and then used to generate data for the same problem. This has the advantage that a GAN creates data in a significantly reduced computing time. N training events for a GAN can result in GN generated events with the gain factor, G, being more than one. This appears to violate the principle that one cannot get information for free. This is not the only way to amplify data so this process will be referred to as data amplification which is studied using information theoretic concepts. It is shown that a gain of greater than one is possible whilst keeping the information content of the data unchanged. This leads to a mathematical bound which only depends on the number of generated and training events. This study determines conditions on both the underlying and reconstructed probability distributions to ensure this bound. In particular, the resolution of variables in amplified data is not improved by the process but the increase in sample size can still improve statistical significance. The bound is confirmed using computer simulation and analysis of GAN generated data from the literature.

生成模型数据放大信息论统计显著性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。