新基准PhyCoBench可自动评估视频生成模型的物理合理性。
A Physical Coherence Benchmark for Evaluating Video Generation Models via Optical Flow-guided Frame Prediction
- 用光流引导帧预测构建自动化评估模型PhyCoPredictor。
- 在120个物理原则提示上测试,人工与自动评估高度一致。
- 适合关注视频真实性的研究人员和开发者。
近期视频生成模型虽展现世界模拟潜力,但常违背物理规律,而多数文本到视频基准未关注此问题。我们提出专门评估视频物理一致性的基准PhyCoBench,包含120个覆盖7类物理原理的提示,捕捉视频中可观测的关键物理法则。我们在PhyCoBench上评估了四种前沿文本到视频模型,并进行人工评估。同时提出自动化评估模型PhyCoPredictor——一个分步生成光流与视频帧的扩散模型。通过对比自动与人工排序的一致性,实验表明PhyCoPredictor当前最接近人类评价,可有效评估视频物理一致性,为未来模型优化提供洞见。基准数据集、提示、评估工具PhyCoPredictor及生成视频已开源至GitHub:https://github.com/Jeckinchen/PhyCoBench。
原文摘要 · Abstract (English)
Recent advances in video generation models demonstrate their potential as world simulators, but they often struggle with videos deviating from physical laws, a key concern overlooked by most text-to-video benchmarks. We introduce a benchmark designed specifically to assess the Physical Coherence of generated videos, PhyCoBench. Our benchmark includes 120 prompts covering 7 categories of physical principles, capturing key physical laws observable in video content. We evaluated four state-of-the-art (SoTA) T2V models on PhyCoBench and conducted manual assessments. Additionally, we propose an automated evaluation model: PhyCoPredictor, a diffusion model that generates optical flow and video frames in a cascade manner. Through a consistency evaluation comparing automated and manual sorting, the experimental results show that PhyCoPredictor currently aligns most closely with human evaluation. Therefore, it can effectively evaluate the physical coherence of videos, providing insights for future model optimization. Our benchmark, including physical coherence prompts, the automatic evaluation tool PhyCoPredictor, and the generated video dataset, has been released on GitHub at https://github.com/Jeckinchen/PhyCoBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。