用关键帧控制扩散模型,模拟人群行人的真实移动轨迹。
Can Image-To-Video Models Simulate Pedestrian Dynamics?
- 以行人轨迹基准数据的关键帧为条件,引导图像到视频模型生成
- 在拥挤公共场景中成功复现了真实行人运动模式
- 适合关注行为建模与视频生成融合的科研人员
基于扩散变换器(DiT)的高性能图像到视频(I2V)模型,在大规模视频数据集上训练后展现出强大的世界建模能力。我们研究这些模型是否能在拥挤公共场所生成真实的行人运动模式。方法上,利用行人轨迹基准数据提取关键帧作为条件,输入I2V模型进行视频生成,并通过行人动力学的定量指标评估其轨迹预测性能。
原文摘要 · Abstract (English)
Recent high-performing image-to-video (I2V) models based on variants of the diffusion transformer (DiT) have displayed remarkable inherent world-modeling capabilities by virtue of training on large scale video datasets. We investigate whether these models can generate realistic pedestrian movement patterns in crowded public scenes. Our framework conditions I2V models on keyframes extracted from pedestrian trajectory benchmarks, then evaluates their trajectory prediction performance using quantitative measures of pedestrian dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。