无需3D重建,用文本控制视角和运镜,实现高质量视频重拍。
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

- 基于时间步敏感性分析,利用高噪声阶段学习相机运动与视觉表征。
- 无需配对数据或3D重建,支持大范围运镜与视角切换,生成新视域内容。
- 适合需要灵活控制镜头视角的视频生成、虚拟拍摄等场景。
视频重拍旨在通过可控的摄像机运动与视角重生成视频。现有方法依赖显式3D先验,受限于重建质量,且在合成未见区域时表现不佳;或依赖不同轨迹的配对视频,但数据稀缺制约泛化能力。本文通过文本驱动的语义视角指定,实现对镜头尺度、视角角度及第一/第三人称视角的控制。提出无需3D结构的TARS框架:时间步敏感性分析表明,相机运动主要在高噪声阶段建立,此时形成粗粒度时空结构。基于此,引入自监督训练,无需配对重拍数据或3D重建,即可学习相机动态与基础视觉表示。结合数据缩放与联合文本-相机条件控制,TARS支持鲁棒的镜头与视角控制,在大幅运镜下仍可合理生成源视图外区域内容,并实现反向视角重拍与视角切换。大量实验表明,相较以往方法,TARS在相机控制精度与时间一致性上均更优。
原文摘要 · Abstract (English)
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。