arXiv:2605.25266cs.CV2026-05

让视频生成模型可精准控制镜头焦距、光圈等成像参数变化。

DeltaCam: Differential Intrinsic Camera Modeling for Video Generation

论文配图:DeltaCam: Differential Intrinsic Camera Modeling for Video Generation
图 1 · 摘自论文原文
  • 用相对变化量建模相机参数,避免依赖精确标签。
  • 在合成数据上训练后,可平滑控制焦距、光圈、色温等参数变化。
  • 适合需要真实感镜头效果的视频生成与编辑场景。

将相机内参融入视频生成模型,能系统性控制场景动态与影响视觉外观的成像过程。以往工作主要关注外参控制(如相机位姿和运动),而将内参视为隐含或固定。核心瓶颈在于缺乏具有准确且多样时变相机元数据的大规模视频数据集,导致学习绝对相机参数化困难。因此,现有模型难以以可控且时间一致的方式引入景深变化、曝光波动、镜头畸变和色彩处理等摄影行为。我们提出DeltaCam,一种基于Δ-参数化神经相机适配器的视频扩散框架,通过相对变化而非绝对状态来建模相机行为。在合成视频数据上学习该微分形式后,降低了对真实世界相机标签的依赖,并实现了对焦距、光圈、ISO、色温及镜头畸变等成像因素的平滑、一致控制。通过两种机制扩展至真实画面:一是在真实图像-元数据对上微调控制以实现精确镜头匹配;二是提取解耦嵌入,实现无需显式相机参数的视频到视频风格迁移。通过有效分离场景内容与成像行为,DeltaCam支持现有模型难以实现的相机一致性视频生成与编辑。最终结果表明,该方法为连接合成控制与真实摄影模拟提供了一种实用且可扩展的路径。

原文摘要 · Abstract (English)

Incorporating camera intrinsics into video generation models offers a principled way to control not only scene dynamics but also the imaging process that governs visual appearance. Prior work has primarily focused on extrinsic control, such as camera pose and motion, while treating intrinsic camera parameters as implicit or fixed. A key bottleneck is the lack of large-scale video datasets with accurate and diverse temporally varying camera metadata, which makes learning absolute camera parameterizations difficult. As a result, current models struggle to incorporate photographic camera behavior, including depth-of-field transitions, exposure variations, lens distortions, and color processing, in a controllable and temporally consistent manner. We introduce DeltaCam, a video diffusion framework that models camera behavior through $Δ$-parameterized neural camera adaptors, operating on relative changes in camera motion and intrinsics instead of absolute states. By learning this differential formulation from synthetic video data, we mitigate reliance on precise real-world camera labels and enable smooth, consistent control over imaging factors such as focal length, aperture, ISO, color temperature, and lens distortion. We extend this framework to real-world footage through two mechanisms: finetuning the controls on real image-metadata pairs for precise shot matching, and extracting disentangled embeddings for implicit video-to-video style transfer without requiring explicit camera parameters. By effectively separating scene content from intrinsic imaging behavior, DeltaCam enables camera-consistent video generation and editing operations that are difficult to achieve with existing models. Ultimately, our results establish a practical and scalable approach for bridging synthetic control and real-world photographic emulation.

视频生成扩散模型相机建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。