用位移场指导视频生成,无需训练就能实现自然相机控制。
Probing into Camera Control of Video Models

- 将相机控制转化为帧间特征的可微位移场
- 相比微调基线,多指标表现接近且损失小
- 可无训练探测模型相机控制能力,适合研究者使用
视频是丰富的3D/4D视觉观测源,相机控制是生成模型产生几何合理内容的关键。现有方法通常依赖额外相机模块和成对数据学习运动到视频的映射,但这类数据规模小、多样性差、场景动态有限,导致模型输出分布狭窄,削弱基础模型的强先验。本文提出新视角:相机控制不必建模为隐式映射,而可作为几何引导,通过在去噪过程中对潜在特征进行可微重采样,施加帧间位移。该方法仅需极少调整,在多个质量指标上优于微调基线,且适用于多数视频扩散模型而无需训练。进一步,该方法可作为探针,揭示主流视频模型共享的普遍偏差及对相机控制响应差异,并在多视角生成任务中进行基准测试,为3D/4D任务提供洞见。
原文摘要 · Abstract (English)
Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometrically meaningful content. Existing approaches typically learn a mapping from camera motion to video using additional camera modules and paired data. However, such datasets are often limited in scale, diversity, and scene dynamics, which can bias the model toward a narrow output distribution and compromise the strong prior learned by the base model. These limitations motivate a different perspective on camera control. In this paper, we show that camera control need not be modeled as an implicit mapping problem, but can instead be treated as a form of geometric guidance that induces displacements across frames. Specifically, we reformulate camera control into a set of displacement fields and apply them via differentiable resampling of latent features during denoising. Our simple approach achieves effective camera control with minimal degradation across diverse quality metrics compared to fine-tuned baselines. Since our method is applicable to most video diffusion models without training, it can also serve as a probe to study the camera control capabilities of base models. Using this probe, we identify universal biases shared by representative video models, as well as disparities in their responses to camera control. Finally, we benchmark their performance in multi-view generation, offering insights into their potential for 3D/4D tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。