让视频生成同时控制相机与物体运动,实现多视角一致输出
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention

- 用视图融合注意力机制增强视角与时间一致性
- 支持多路径相机轨迹生成,且保持物体运动自然
- 适合需要精准相机控制的影视/游戏场景生成
近年来图像到视频生成取得了显著进展,但生成帧的三维一致性与相机可控性仍待解决。现有研究虽尝试引入相机控制,但结果多局限于简单轨迹,或无法为同一场景生成多条不同相机路径的连续视频。为此,我们提出Cavia,一种新型相机可控、多视角视频生成框架,可将输入图像转换为多个时空一致的视频。该框架将空间和时间注意力模块扩展为视图融合注意力模块,提升视角与时间一致性。灵活设计支持联合训练多种数据源:场景级静态视频、对象级合成多视角动态视频、真实世界单目动态视频。据我们所知,Cavia是首个在精确指定相机运动的同时,能生成自然物体运动的模型。大量实验表明,Cavia在几何一致性和感知质量上优于现有方法。
原文摘要 · Abstract (English)
In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are often limited to simple trajectories or lack the ability to generate consistent videos from multiple distinct camera paths for the same scene. To address these limitations, we introduce Cavia, a novel framework for camera-controllable, multi-view video generation, capable of converting an input image into multiple spatiotemporally consistent videos. Our framework extends the spatial and temporal attention modules into view-integrated attention modules, improving both viewpoint and temporal consistency. This flexible design allows for joint training with diverse curated data sources, including scene-level static videos, object-level synthetic multi-view dynamic videos, and real-world monocular dynamic videos. To our best knowledge, Cavia is the first of its kind that allows the user to precisely specify camera motion while obtaining object motion. Extensive experiments demonstrate that Cavia surpasses state-of-the-art methods in terms of geometric consistency and perceptual quality. Project Page: https://ir1d.github.io/Cavia/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。