FVD评估视频生成有缺陷,新指标JEDi更准更快。
Beyond FVD: Enhanced Evaluation Metrics for Video Generation Quality
- 用JEPA模型提取特征,通过多项式MMD计算距离
- 只需FVD16%样本量就稳定,人类评分匹配度提升34%
- 适合视频生成研究者做质量评估
Fréchet Video Distance (FVD) 是广泛使用的视频生成分布质量评估指标,但其有效性依赖于若干关键假设。我们分析发现三大局限:(1) Inflated 3D Convnet (I3D) 特征空间非高斯;(2) I3D 特征对时间扭曲不敏感;(3) 可靠估计需极大量样本。这些削弱了 FVD 的可靠性,表明其不适合作为视频生成评估的唯一指标。经过对多种度量和骨干架构的广泛分析,我们提出 JEDi(JEPA Embedding Distance),基于联合嵌入预测架构(Joint Embedding Predictive Architecture)的特征,采用带多项式核的 Maximum Mean Discrepancy 进行衡量。在多个开源数据集上的实验表明,JEDi 显著优于广泛使用的 FVD,仅需 16% 样本即达稳定值,且平均提升与人工评价的一致性 34%。
原文摘要 · Abstract (English)
The Fréchet Video Distance (FVD) is a widely adopted metric for evaluating video generation distribution quality. However, its effectiveness relies on critical assumptions. Our analysis reveals three significant limitations: (1) the non-Gaussianity of the Inflated 3D Convnet (I3D) feature space; (2) the insensitivity of I3D features to temporal distortions; (3) the impractical sample sizes required for reliable estimation. These findings undermine FVD's reliability and show that FVD falls short as a standalone metric for video generation evaluation. After extensive analysis of a wide range of metrics and backbone architectures, we propose JEDi, the JEPA Embedding Distance, based on features derived from a Joint Embedding Predictive Architecture, measured using Maximum Mean Discrepancy with polynomial kernel. Our experiments on multiple open-source datasets show clear evidence that it is a superior alternative to the widely used FVD metric, requiring only 16% of the samples to reach its steady value, while increasing alignment with human evaluation by 34%, on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。