arXiv:2510.00806cs.CV2025-10被引 2

用视觉语言模型预测物理合理轨迹,指导视频生成

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

  • 先用视觉语言模型预测符合物理规律的粗粒度运动轨迹
  • 在UCF-101和MSR-VTT上实现545和539的FVD得分
  • 适合关注视频生成物理一致性的研究者

当前视频生成模型常产生违背现实物理规律的运动。我们提出TrajVLM-Gen,一种两阶段的物理感知图像到视频生成框架。首先,利用视觉语言模型预测保持真实世界物理一致性的粗粒度运动轨迹;其次,通过基于注意力的机制,以这些轨迹为指导,对细粒度运动进行精修。我们基于视频跟踪数据构建了一个包含真实运动模式的轨迹预测数据集。在UCF-101和MSR-VTT上的实验表明,TrajVLM-Gen优于现有方法,在这两个数据集上分别取得545和539的FVD得分。

原文摘要 · Abstract (English)

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Vision Language Model to predict coarse-grained motion trajectories that maintain consistency with real-world physics. Second, these trajectories guide video generation through attention-based mechanisms for fine-grained motion refinement. We build a trajectory prediction dataset based on video tracking data with realistic motion patterns. Experiments on UCF-101 and MSR-VTT demonstrate that TrajVLM-Gen outperforms existing methods, achieving competitive FVD scores of 545 on UCF-101 and 539 on MSR-VTT.

视频生成轨迹预测扩散模型物理一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。