arXiv:2603.19048cs.CV2026-03中稿 · ECCV

提出新指标SGC,精准衡量动态视频的3D空间几何一致性。

Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation

  • 通过多区域相机位姿差异度量几何一致性
  • 在真实与生成视频上均能有效识别几何错误
  • 适合评估生成视频中背景结构的稳定性

近期生成模型可产出高保真视频,但常存在3D空间几何不一致问题。现有评估方法难以准确刻画此类问题:以FVD为代表的保真度指标对几何失真不敏感,而聚焦一致性的基准又会误罚有效的前景运动。为此,我们提出SGC,用于评估动态生成视频中的3D空间几何一致性。该方法通过测量从不同局部区域估计出的多个相机位姿之间的发散程度来量化几何一致性。首先分离静态与动态区域,再将静态背景划分为空间连贯的子区域;对每个像素预测深度,为每个子区域估计局部相机位姿,并计算这些位姿间的发散值以表征几何一致性。在真实视频和生成视频上的实验表明,SGC能稳健量化几何不一致,有效识别出现有指标遗漏的关键失败案例。

原文摘要 · Abstract (English)

Recent generative models can produce high-fidelity videos, yet they often exhibit 3D spatial geometric inconsistencies. Existing evaluation methods fail to accurately characterize these inconsistencies: fidelity-centric metrics like FVD are insensitive to geometric distortions, while consistency-focused benchmarks often penalize valid foreground dynamics. To address this gap, we introduce SGC, a metric for evaluating 3D \textbf{S}patial \textbf{G}eometric \textbf{C}onsistency in dynamically generated videos. We quantify geometric consistency by measuring the divergence among multiple camera poses estimated from distinct local regions. Our approach first separates static from dynamic regions, then partitions the static background into spatially coherent sub-regions. We predict depth for each pixel, estimate a local camera pose for each subregion, and compute the divergence among these poses to quantify geometric consistency. Experiments on real and generative videos demonstrate that SGC robustly quantifies geometric inconsistencies, effectively identifying critical failures missed by existing metrics.

视频生成几何一致性评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。