arXiv:2412.01429cs.CV2024-12被引 12

让视频生成精准跟随镜头运动轨迹,提升长视频一致性

CPA: Camera-pose-awareness Diffusion Transformer for Video Generation

论文配图:CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
图 1 · 摘自论文原文
  • 用稀疏运动编码模块将镜头位姿转为时空嵌入
  • 注入运动块使镜头与物体运动更连贯,长视频表现更优
  • 适合需要精确镜头控制的影视、动画生成场景

尽管基于扩散变换器(DiT)的方法在视频生成上取得显著进展,但在可控镜头视角方面仍存在明显差距。现有方法如OpenSora未能精确遵循预期运动轨迹和物理交互,限制了下游应用灵活性。为此,我们提出CPA,一种统一的文本到视频生成方法,具备镜头位姿感知能力,能同时建模文本、视觉与空间条件。具体地,我们设计稀疏运动编码(SME)模块,将镜头位姿信息转化为时空嵌入,并通过时间注意力注入(TAI)模块将运动片段注入每个时空扩散变换器(ST-DiT)块。该插件式架构兼容原生DiT参数,支持多种镜头姿态与灵活物体运动。大量定性与定量实验表明,本方法在长视频生成中优于基于LDM的方法,在轨迹一致性和物体一致性方面达到最优性能。

原文摘要 · Abstract (English)

Despite the significant advancements made by Diffusion Transformer (DiT)-based methods in video generation, there remains a notable gap with controllable camera pose perspectives. Existing works such as OpenSora do NOT adhere precisely to anticipated trajectories and physical interactions, thereby limiting the flexibility in downstream applications. To alleviate this issue, we introduce CPA, a unified camera-pose-awareness text-to-video generation approach that elaborates the camera movement and integrates the textual, visual, and spatial conditions. Specifically, we deploy the Sparse Motion Encoding (SME) module to transform camera pose information into a spatial-temporal embedding and activate the Temporal Attention Injection (TAI) module to inject motion patches into each ST-DiT block. Our plug-in architecture accommodates the original DiT parameters, facilitating diverse types of camera poses and flexible object movement. Extensive qualitative and quantitative experiments demonstrate that our method outperforms LDM-based methods for long video generation while achieving optimal performance in trajectory consistency and object consistency.

视频生成扩散模型镜头控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。