用点轨迹建模运动,更准确评估生成视频的时序一致性。
Direct Motion Models for Assessing Generated Videos
- 基于点轨迹自编码,捕捉视频中物体的运动特征。
- 在生成视频的时序真实感评估上优于FVD等主流指标。
- 可定位生成视频中的时空异常,提升错误可解释性。
当前生成视频模型虽能生成逼真帧,但运动质量差,而现有评价方法如FVD难以捕捉此问题。本文提出一种新度量方法,基于点轨迹的自编码器生成运动特征,不仅能比较生成视频与真实视频的分布(最少一对,或多对数据集),还可单独评估单个视频的运动质量。相比像素重建或动作识别特征,该方法对合成数据中的时间扭曲更敏感,能更好预测人类对生成视频时序一致性和真实感的评价。此外,通过点轨迹表示,可实现时空定位生成视频中的不一致区域,提供比以往方法更强的错误可解释性。项目主页含结果概览及代码链接:http://trajan-paper.github.io。
原文摘要 · Abstract (English)
A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond FVD by developing a metric which better measures plausible object interactions and motion. Our novel approach is based on auto-encoding point tracks and yields motion features that can be used to not only compare distributions of videos (as few as one generated and one ground truth, or as many as two datasets), but also for evaluating motion of single videos. We show that using point tracks instead of pixel reconstruction or action recognition features results in a metric which is markedly more sensitive to temporal distortions in synthetic data, and can predict human evaluations of temporal consistency and realism in generated videos obtained from open-source models better than a wide range of alternatives. We also show that by using a point track representation, we can spatiotemporally localize generative video inconsistencies, providing extra interpretability of generated video errors relative to prior work. An overview of the results and link to the code can be found on the project page: http://trajan-paper.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。