让视频生成同时控制光影与几何,实现更真实的光照变化。
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
- 融合HDR光照图、重光照帧和3D点轨迹,统一建模光影与结构。
- 支持文本或背景条件下的可控重光照,生成视频时序一致。
- 相比现有方法,光照还原更真实,适合影视特效与虚拟场景创作。
尽管基于扩散模型的视频生成方法能从文本或图像输入生成高质量、高分辨率视频序列,但在控制场景光照与跨帧视觉表现时缺乏显式的几何线索整合。为此,我们提出IllumiCraft,一种端到端的扩散框架,接受三种互补输入:(1) 高动态范围(HDR)视频图以实现精细光照控制;(2) 经合成重光照的帧(可选配静态背景参考图)提供外观线索;(3) 3D点轨迹捕捉精确的三维几何信息。通过在统一扩散架构中融合光照、外观与几何线索,IllumiCraft生成与用户提示对齐的时序连贯视频。该方法支持背景条件和文本条件下的视频重光照,并在保真度上优于现有可控视频生成方法。
原文摘要 · Abstract (English)
Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appearance across frames. To address this limitation, we propose IllumiCraft, an end-to-end diffusion framework accepting three complementary inputs: (1) high-dynamic-range (HDR) video maps for detailed lighting control; (2) synthetically relit frames with randomized illumination changes (optionally paired with a static background reference image) to provide appearance cues; and (3) 3D point tracks that capture precise 3D geometry information. By integrating the lighting, appearance, and geometry cues within a unified diffusion architecture, IllumiCraft generates temporally coherent videos aligned with user-defined prompts. It supports background-conditioned and text-conditioned video relighting and provides better fidelity than existing controllable video generation methods. Project Page: https://yuanze-lin.me/IllumiCraft_page
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。