无需参考视频,通过视觉里程计与光流分析评估生成视频的物理一致性。
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

- 结合相对与绝对评估,利用DROID-SLAM和SEA-RAFT检测物理不一致
- 相对评估使任务成功率提升超8%,缩小仿真与现实差距
- 可定位物理错误发生的时间与空间位置,适合视频生成质量验证
我们提出无参考的物理一致性评估方法,结合相对与绝对评估以衡量生成视频的保真度。尽管WorldGym或WorldEval等工具可通过视频生成实现机器人仿真,但物理保真度不足常导致虚拟环境无法准确复现真实世界中VLA模型的任务成功率。不同于依赖人工评分(Elo)或不可用的真值参考(FVD)的现有方法,本工作利用DROID-SLAM和SEA-RAFT量化物理不一致性,受WorldScore启发。经相对一致性评估筛选后的视频,任务成功率提升超过8%,有效缩小了仿真到现实的差距。此外,绝对评估支持时空定位,可可视化物理伪影出现的具体时间与位置。
原文摘要 · Abstract (English)
We introduce reference-free measures for evaluating the physical consistency of generated videos, combining relative and absolute approaches to assess fidelity. Although tools like WorldGym or WorldEval enable robotic simulation via video generation, physical fidelity gaps often prevent these environments from accurately reproducing real-world task success rates of VLA models. Unlike existing evaluation methods, which require costly human voting (Elo) or unavailable ground-truth references (FVD), our approach utilizes DROID-SLAM and SEA-RAFT to quantify physical inconsistencies, motivated by WorldScore. Videos filtered using our relative consistency assessment show an improvement in task success rates of over 8%, effectively narrowing the simulation-to-reality gap. Furthermore, our absolute assessment enables spatio-temporal localization, providing visualization of when and where physical artifacts occur.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。