用3D重建评估文本生成视频的场景一致性,发现现有模型常存在多视角不一致问题。
GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction
- 通过相机位姿估计与静态3D代理重建,量化视频生成的3D一致性
- 在80个静态场景上完成3840次重建,发现运动、渲染误差等指标常矛盾
- 适合关注生成视频几何真实性的研究人员和开发者
基于相机提示的文本生成视频(T2V)模型广泛用于合成虚拟摄像机拍摄内容,如环绕物体或穿越静态场景。仅保证视觉合理性不足,生成帧需能支持对单一静态3D场景的连贯多视角重建。我们提出GeoT2V-Bench,一种基于重建的诊断基准,用于评估相机提示的T2V片段是否支持刚性3D重建。该方法使用VGGT风格几何估计获取每帧相机内参与位姿,拟合DeformableGS,通过时间中值聚合生成静态MedianGS代理,并沿估计相机路径渲染。不同于简单通过/失败或单值评分,GeoT2V-Bench输出包含图像运动明显度、轨迹行为、静态渲染误差、静态渲染光流一致性及柔性与静态拟合差距的连续重建轮廓。在12个开源模型配置、80个GeCo-Eval静态场景提示、四种子采样下的公平评估中,共完成3840次重建,结果显示可见运动、静态渲染误差、光流一致性与柔性-静态拟合差异常不一致,揭示出将生成视频视为全局静态场景采集时浮现的互补失败模式。
原文摘要 · Abstract (English)
Camera-prompted text-to-video (T2V) models are increasingly used to synthesize virtual camera captures, such as orbiting objects or moving through static scenes. For these outputs, visual plausibility is insufficient: the generated frames should also provide coherent multi-view evidence for a single static 3D scene. We introduce GeoT2V-Bench, a reconstruction-based diagnostic benchmark for evaluating whether camera-prompted T2V clips can support explicit rigid 3D reconstruction. Our pipeline estimates per-frame camera intrinsics and poses with VGGT-style geometry estimation, fits DeformableGS, derives a static MedianGS proxy by temporal-median aggregation, and renders this proxy along the estimated camera path. Instead of producing a pass/fail label or a single scalar score, GeoT2V-Bench reports a continuous reconstruction profile covering apparent image motion, estimated trajectory behavior, MedianGS static rendering error, static-render flow agreement, and the gap between flexible and static fits. On a fair-format four-seed evaluation with 3,840 completed reconstructions from 12 open-weight model configurations and 80 GeCo-Eval static-scene prompts, we find that visible motion, static rendering error, flow agreement, and flexible-vs-static behavior often disagree. GeoT2V-Bench therefore captures complementary failure modes that emerge when generated videos are tested as global static-scene acquisitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。