arXiv:2412.14340cs.LGcs.AI2024-12AAAI被引 1

用信息论统一评估生成模型的保真度与多样性。

A Unifying Information-theoretic Perspective on Evaluating Generative Models

  • 基于kNN密度估计,从信息论视角统一多种评估指标。
  • 提出三维指标:保真度、类间/类内多样性分别量化。
  • 适用于任意数据集,可对样本和模式级进行分析。

鉴于生成模型输出难以解释,当前研究聚焦于构建有意义的评估指标。近年方法借用分类领域的“精确率”与“召回率”,分别衡量输出保真度(真实感)和多样性(真实数据变化的覆盖)。随着指标数量增加,亟需统一视角以促进比较与理解优劣。本文基于k近邻(kNN)密度估计,将一类kNN-based指标统一于信息论框架下。同时提出三维指标:精确率交叉熵(PCE)、召回率交叉熵(RCE)与召回率熵(RE),分别衡量保真度及两类多样性——类间与类内。该无领域依赖指标源自熵与交叉熵的信息论概念,支持样本级与模式级分解分析。详尽实验表明,各分量对对应质量敏感,并揭示了其他指标的不良行为。

原文摘要 · Abstract (English)

Considering the difficulty of interpreting generative model output, there is significant current research focused on determining meaningful evaluation metrics. Several recent approaches utilize "precision" and "recall," borrowed from the classification domain, to individually quantify the output fidelity (realism) and output diversity (representation of the real data variation), respectively. With the increase in metric proposals, there is a need for a unifying perspective, allowing for easier comparison and clearer explanation of their benefits and drawbacks. To this end, we unify a class of kth-nearest-neighbors (kNN)-based metrics under an information-theoretic lens using approaches from kNN density estimation. Additionally, we propose a tri-dimensional metric composed of Precision Cross-Entropy (PCE), Recall Cross-Entropy (RCE), and Recall Entropy (RE), which separately measure fidelity and two distinct aspects of diversity, inter- and intra-class. Our domain-agnostic metric, derived from the information-theoretic concepts of entropy and cross-entropy, can be dissected for both sample- and mode-level analysis. Our detailed experimental results demonstrate the sensitivity of our metric components to their respective qualities and reveal undesirable behaviors of other metrics.

生成模型评估指标信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。