arXiv:2602.21390cs.LGstat.ML2026-02被引 1

提出高效生成难以被验证的多类型数据模型的新方法。

Defensive Generation

  • 基于再生核希尔伯特空间的在线校准,设计高效算法。
  • 生成误差率逼近最优的 T^{-1/2},且支持无限类测试。
  • 适用于需要防伪生成的高维数据场景,如金融、医疗建模。

我们研究如何在在线场景下高效生成标量、多分类和向量值结果的生成模型,使其无法被观测数据和预设计算测试所证伪。贡献有二:首先,我们将在线高维多校准与再生核希尔伯特空间(RKHS)的最新进展结合,实现了前者的高效算法;其次,将该算法应用于结果不可区分性问题。提出的 Defensive Generation 方法是首个能高效生成非伯努利型结果的在线不可区分生成模型,且对无限类测试(包括高阶矩检验)具有不可证伪性。该方法在样本数上近线性运行,生成误差达到最优的渐近消失率 T^{-1/2}。

原文摘要 · Abstract (English)

We study the problem of efficiently producing, in an online fashion, generative models of scalar, multiclass, and vector-valued outcomes that cannot be falsified on the basis of the observed data and a pre-specified collection of computational tests. Our contributions are twofold. First, we expand on connections between online high-dimensional multicalibration with respect to an RKHS and recent advances in expected variational inequality problems, enabling efficient algorithms for the former. We then apply this algorithmic machinery to the problem of outcome indistinguishability. Our procedure, Defensive Generation, is the first to efficiently produce online outcome indistinguishable generative models of non-Bernoulli outcomes that are unfalsifiable with respect to infinite classes of tests, including those that examine higher-order moments of the generated distributions. Furthermore, our method runs in near-linear time in the number of samples and achieves the optimal, vanishing T^{-1/2} rate for generation error.

生成模型在线学习防伪生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。