评估视频生成模型在行人动态上的真实感,发现其行为合理但存在穿模问题。
PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation
- 用无相机参数的2D俯视轨迹重建法评估多行人交互
- 领先模型在人群密度和互动类型上表现合理,但有行人消失/合并现象
- 适合研究视频生成与物理一致性、人机交互的学者参考
行人模拟传统依赖人工调参的规则模型,难以扩展。而大规模视频生成模型虽在视觉上逼真,但现有评估仅关注单人生成,未检验多人交互场景的真实性。本文提出严谨评估协议,将文本到视频(T2V)和图像到视频(I2V)模型作为隐式行人动力学模拟器进行评测。针对I2V,使用已有数据集的起始帧实现与真实视频的直接对比;针对T2V,设计涵盖不同人群密度与互动类型的提示套件。核心是无需已知相机参数,从像素空间重建2D鸟瞰轨迹的方法。分析表明,主流模型具备合理的多人行为先验,但存在行人合并与消失等问题,暴露出物理一致性的局限。
原文摘要 · Abstract (English)
Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale video generation models have achieved high visual realism across diverse settings, motivating exploration of their potential as general-purpose world simulators. Existing benchmarks primarily assess single-subject realism rather than scenes with multiple interacting people, leaving the plausibility of multi-agent dynamics in generated videos untested. We propose a rigorous evaluation protocol to benchmark text-to-video (T2V) and image-to-video (I2V) models as implicit simulators of pedestrian dynamics. For I2V, we leverage start frames from established datasets to enable direct comparison with ground truth videos, while for T2V we design a prompt suite covering varied crowd densities and interaction types. A key component is a method to reconstruct 2D bird's-eye view trajectories from pixel-space without known camera parameters. Our analysis shows that leading models exhibit effective priors for plausible multi-agent behavior, though issues such as merging and disappearing pedestrians reveal limits to their physical consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。