arXiv:2502.07531cs.CVcs.AI2025-02中稿 · TVCG 2026被引 25

实现图像到视频生成中相机、物体和光照的精准协同控制

VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation

  • 通过跨因素交互建模,统一控制相机、物体与光照运动
  • 在多个场景下实现最高精度的控制与视觉一致性
  • 适合需要精细动态控制的影视创作与虚拟仿真用户

可控图像到视频(I2V)生成将参考图像转化为连贯视频,由用户指定控制信号引导。精确控制相机运动、物体运动和光照方向对高质量生成至关重要,但现有方法常独立处理这些因素,忽视了真实场景中视角、几何与光照之间的物理耦合,导致阴影错位、透视漂移等视觉不一致问题。我们提出VidCRAFT3,一个统一且灵活的I2V框架,显式建模几何、运动与光照间的跨因子交互,支持相机、物体和光照的独立与联合控制。Image2Cloud提供显式的3D几何先验以实现精准相机运动控制;ObjMotionNet将稀疏物体轨迹编码为多尺度运动特征,指导真实物体运动;Spatial Triple-Attention Transformer通过光照交叉注意力集成光照方向,实现一致的再打光。为解决联合标注数据稀缺问题,我们构建了带有逐帧光照方向标注的VideoLightingDirection(VLD)数据集,并引入三阶段渐进式训练策略,在无需完全联合标注的情况下实现鲁棒学习。大量实验表明,VidCRAFT3在多种场景下均达到最先进的控制精度与视觉连贯性。

原文摘要 · Abstract (English)

Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. While precise control over camera motion, object motion, and lighting is essential for high-fidelity creation, existing methods often treat these factors independently. This overlooks the physical coupling among viewpoint, geometry, and illumination in dynamic scenes, leading to visual inconsistencies such as mismatched shadows and perspective drift under simultaneous changes. We present VidCRAFT3, a unified and flexible I2V framework that explicitly models cross-factor interactions among geometry, motion, and illumination, enabling both independent and joint control over camera motion, object motion, and lighting direction. Image2Cloud provides explicit 3D geometric priors for accurate camera motion control. ObjMotionNet encodes sparse object trajectories into multi-scale motion features to guide realistic object motion. A Spatial Triple-Attention Transformer integrates lighting direction through lighting cross-attention for consistent relighting. To address the scarcity of jointly annotated data, we construct the VideoLightingDirection (VLD) dataset with accurate per-frame lighting direction annotations, and introduce a three-stage progressive training strategy that enables robust learning without fully joint annotations. Extensive experiments demonstrate that VidCRAFT3 achieves state-of-the-art performance in control precision and visual coherence across diverse scenarios.

图像到视频可控生成光照控制三维几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。