用离散视觉符号统计评估生成模型,更贴近人眼对质量的感知。
Evaluating Generative Models via One-Dimensional Code Distributions
- 在离散视觉符号空间中分析生成图像的质量特征
- 新指标在多个数据集上与人类判断相关性达领先水平
- 适合关注生成图像感知质量的研究者和评测人员
现有生成模型评估多依赖于连续特征分布指标(如FID),这些指标忽略影响感知质量的关键视觉线索。本文转而考察离散视觉标记空间,利用现代一维图像标记器同时编码语义与感知信息,使质量体现为可预测的标记统计规律。提出无需训练的代码本直方图距离(CHD)和基于合成退化的无参考质量评分(CMMS)。为全面测试指标鲁棒性,构建了包含21万张图像、62种视觉形态和12种生成模型的VisForm基准,由专家标注。在AGIQA、HPDv2/3和VisForm上,所提方法在与人类判断的相关性上达到当前最优。代码与数据集将公开,地址:https://github.com/zexiJia/1d-Distance。
原文摘要 · Abstract (English)
Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training-free distribution metric in token space, and Code Mixture Model Score (CMMS), a no-reference quality metric learned from synthetic degradations of token sequences. To stress-test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token-based metrics achieve state-of-the-art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d-Distance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。