提出无需假设的图像质量评估方法,突破传统指标局限。
APEX: Assumption-free Projection-based Embedding eXamination Metric for Image Quality Assessment

- 用切片瓦瑟斯坦距离构建无假设的相似性度量,避免参数化偏差。
- 在多种图像退化下表现更稳健,跨数据集评估稳定性强。
- 兼容CLIP、DINOv2等开放词汇模型,适用于各类生成图像评估。
随着生成模型视觉质量不断提升,图像评估仍依赖传统的特征分布指标(如FID)。然而,这些指标受限于过时特征的封闭词表瓶颈和刚性参数化假设。近期方法虽采用现代主干网络缓解特征瓶颈,但仍受参数化限制。为此,我们提出APEX(Assumption-free Projection-based Embedding eXamination),一种基于切片瓦瑟斯坦距离的新评估框架,该距离具有数学基础且无需假设。理论与实证证明其在高维空间中具备良好可扩展性。APEX不依赖特定嵌入,使用CLIP和DINOv2两个开放词汇基础模型作为特征提取器。基准测试显示,相较于现有基线,APEX对视觉退化更具鲁棒性,且在同域与跨域数据集上均表现出优异的稳定性。
原文摘要 · Abstract (English)
As generative models achieve unprecedented visual quality, the gold standard for image evaluation remains traditional feature-distribution metrics (e.g., FID). However, these metrics are provably hindered by the closed-vocabulary bottleneck of outdated features and the assumptive bias of rigid parametric formulations. Recent alternatives exploit modern backbones to solve the feature bottleneck, yet continue to suffer from parametric limitations. To close this gap, we introduce APEX (Assumption-free Projection-based Embedding eXamination), a novel evaluation framework leveraging the Sliced Wasserstein Distance as a mathematically grounded, assumption-free similarity measure. APEX inherits effective scalability to high-dimensional spaces, as we prove with theoretical and empirical evidences. Moreover, APEX is embedding-agnostic and uses two open-vocabulary foundation models, CLIP and DINOv2, as feature extractors. Benchmarking APEX against established baselines reveals superior robustness to visual degradations. Additionally, we show that APEX metrics exhibit intra- and cross-dataset stability, ensuring highly stable evaluations on out-of-domain datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。