提出量化评估视频生成几何一致性的新方法,可发现传统指标忽略的结构错误。
Quantitative Video World Model Evaluation for Geometric-Consistency

- 通过单目重建与点追踪获取物体三维坐标,计算投影几何残差
- 在多个主流视频生成模型上发现一致性几何失效模式
- 适合关注物理真实性的视频生成研究者使用
生成式视频模型日益被视为隐式世界模型,但评估其是否生成符合物理规律的三维结构与运动仍具挑战。现有评估多依赖人工判断或学习型评分器,主观性强且对几何错误诊断能力弱。本文提出PDI-Bench(视角畸变指数)框架,用于定量审计生成视频的几何一致性。给定生成视频片段,利用分割与点跟踪技术(如SAM 2、MegaSaM、CoTracker3)获取物体中心观测,通过单目重建将其提升至三维世界坐标,计算一组投影几何残差,涵盖尺度-深度对齐、三维运动一致性、三维结构刚性三个失效维度。为支持系统评估,构建PDI-Dataset,覆盖多种施加几何约束的场景。在多个先进视频生成器中,PDI揭示了感知指标未捕捉的几何特异性失效模式,并为迈向物理基础的视频生成与世界模型提供诊断信号。代码与数据集详见https://pdi-bench.github.io/。
原文摘要 · Abstract (English)
Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human judgment or learned graders, which can be subjective and weakly diagnostic for geometric failures. We introduce PDI-Bench (Perspective Distortion Index), a quantitative framework for auditing geometric coherence in generated videos. Given a generated clip, we obtain object-centric observations via segmentation and point tracking (e.g., SAM 2, MegaSaM, and CoTracker3), lift them to 3D world-space coordinates via monocular reconstruction, and compute a set of projective-geometry residuals capturing three failure dimensions: scale-depth alignment, 3D motion consistency, and 3D structural rigidity. To support systematic evaluation, we build PDI-Dataset, covering diverse scenarios designed to stress these geometric constraints. Across state-of-the-art video generators, PDI reveals consistent geometry-specific failure modes that are not captured by common perceptual metrics, and provides a diagnostic signal for progress toward physically grounded video generation and physical world model. Our code and dataset can be found at https://pdi-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。