arXiv:2604.05324cs.LGcs.IT2026-04被引 1

提出生成模型统计可评估性理论框架,厘清哪些指标能可靠评估。

A Theoretical Framework for Statistical Evaluability of Generative Models

  • 构建基于有限样本的生成模型评估理论框架
  • 有界测试类下的IPM可有限样本近似,低复杂度类可任意精度评估
  • 揭示熵类指标因罕见事件不可靠,适合理论研究者参考

统计评估旨在使用从真实分布中独立同分布采样的测试数据来估计模型的泛化性能。在分类等监督学习场景中,误差率等指标定义明确,足够大的数据集下测试误差可可靠逼近总体误差。然而,由于生成模型具有开放性,其评估更具挑战:合适的评估指标不明确,且有限样本下的评估可靠性存疑。本文提出一个生成模型评估的理论框架,建立了常用指标的可评估性结果。研究两类指标:基于测试的指标(如积分概率度量IPMs)和Rényi散度。证明对任意有界测试类,IPMs可在有限样本下实现乘法与加法误差近似;若测试类的fat-shattering维数有限,则可任意精度评估。相反,Rényi和KL散度无法从有限样本可靠评估,因其值可能由罕见事件决定。还分析了困惑度作为评估方法的潜力与局限。

原文摘要 · Abstract (English)

Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d. test data sampled from the ground-truth distribution. In supervised learning settings such as classification, performance metrics such as error rate are well-defined, and test error reliably approximates population error given sufficiently large datasets. In contrast, evaluation is more challenging for generative models due to their open-ended nature: it is unclear which metrics are appropriate and whether such metrics can be reliably evaluated from finite samples. In this work, we introduce a theoretical framework for evaluating generative models and establish evaluability results for commonly used metrics. We study two categories of metrics: test-based metrics, including integral probability metrics (IPMs), and Rényi divergences. We show that IPMs with respect to any bounded test class can be evaluated from finite samples up to multiplicative and additive approximation errors. Moreover, when the test class has finite fat-shattering dimension, IPMs can be evaluated with arbitrary precision. In contrast, Rényi and KL divergences are not evaluable from finite samples, as their values can be critically determined by rare events. We also analyze the potential and limitations of perplexity as an evaluation method.

生成模型评估理论信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。