arXiv:2506.17220cs.CV2025-06NeurIPS被引 27

首次量化分析视频扩散模型如何建立帧间对应关系

Emergent Temporal Correspondences from Video Diffusion Transformers

论文配图:Emergent Temporal Correspondences from Video Diffusion Transformers
图 1 · 摘自论文原文
  • 构建带伪真值追踪标注的数据集,评估3D注意力各组件的作用
  • 发现特定层的查询键相似性在去噪过程中主导帧间匹配
  • 可零样本追踪点,提升生成视频时序一致性,无需额外训练

基于扩散Transformer(DiT)的视频扩散模型在生成时序连贯视频方面取得了显著进展。然而,一个根本性问题仍未解决:这些模型如何在内部建立并表示跨帧的时间对应关系?我们提出DiffTrack,首个定量分析框架,旨在回答这一问题。DiffTrack构建了一个由提示生成的视频数据集,带有伪真值追踪标注,并提出了新颖的评估指标,系统分析DiT中3D注意力机制各组件(如表示、层、时间步)对建立时间对应关系的贡献。分析发现,特定而非所有层中的查询-键相似性在时间匹配中起关键作用,且该匹配在去噪过程中逐渐增强。我们展示了DiffTrack在零样本点追踪中的实际应用,其表现优于现有视觉基础与自监督视频模型。进一步地,我们基于发现提出一种新型引导方法,实现运动增强的视频生成,显著提升生成视频的时序一致性,且无需额外训练。我们认为本工作揭示了视频DiT内部运作机制,为后续研究与应用奠定了基础。

原文摘要 · Abstract (English)

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish and represent temporal correspondences across frames? We introduce DiffTrack, the first quantitative analysis framework designed to answer this question. DiffTrack constructs a dataset of prompt-generated video with pseudo ground-truth tracking annotations and proposes novel evaluation metrics to systematically analyze how each component within the full 3D attention mechanism of DiTs (e.g., representations, layers, and timesteps) contributes to establishing temporal correspondences. Our analysis reveals that query-key similarities in specific, but not all, layers play a critical role in temporal matching, and that this matching becomes increasingly prominent during the denoising process. We demonstrate practical applications of DiffTrack in zero-shot point tracking, where it achieves state-of-the-art performance compared to existing vision foundation and self-supervised video models. Further, we extend our findings to motion-enhanced video generation with a novel guidance method that improves temporal consistency of generated videos without additional training. We believe our work offers crucial insights into the inner workings of video DiTs and establishes a foundation for further research and applications leveraging their temporal understanding.

视频生成扩散模型时序对应注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。