arXiv:2604.07966cs.CV2026-04中稿 · CVPR被引 1

让视频生成精准控制光影与场景,支持自由编辑3D环境。

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

论文配图:Lighting-grounded Video Generation with Renderer-based Agent Reasoning
图 1 · 摘自论文原文
  • 通过3D统一表示解耦光照、布局、镜头轨迹等场景因素
  • 在真实感和时序一致性上达到当前最佳,支持可编辑3D场景
  • 内置智能代理自动将用户指令转为3D控制信号,适合影视制作

扩散模型在视频生成中取得显著进展,但可控性仍是主要瓶颈。场景关键因素如布局、光照、相机轨迹常被纠缠或弱建模,限制其在电影制作与虚拟生产等需精确场景控制领域的应用。本文提出LiVER,一种基于扩散模型的可控视频生成框架。该框架通过新构建的大规模数据集(包含密集标注的对象布局、光照与相机参数)实现对显式3D场景属性的条件生成。方法通过统一3D表示渲染控制信号,解耦各类场景因素。设计轻量级条件模块与渐进训练策略,将信号融入基础视频扩散模型,确保稳定收敛与高保真度。框架支持图像到视频、视频到视频合成,且底层3D场景可完全编辑。进一步开发场景代理,可自动将高层用户指令转化为所需3D控制信号。实验表明,LiVER在保真度与时间一致性上达到最先进水平,同时实现对场景因素的精确、解耦控制,树立了可控视频生成的新标准。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and virtual production where explicit scene control is essential. We present LiVER, a diffusion-based framework for scene-controllable video generation. To achieve this, we introduce a novel framework that conditions video synthesis on explicit 3D scene properties, supported by a new large-scale dataset with dense annotations of object layout, lighting, and camera parameters. Our method disentangles these properties by rendering control signals from a unified 3D representation. We propose a lightweight conditioning module and a progressive training strategy to integrate these signals into a foundational video diffusion model, ensuring stable convergence and high fidelity. Our framework enables a wide range of applications, including image-to-video and video-to-video synthesis where the underlying 3D scene is fully editable. To further enhance usability, we develop a scene agent that automatically translates high-level user instructions into the required 3D control signals. Experiments show that LiVER achieves state-of-the-art photorealism and temporal consistency while enabling precise, disentangled control over scene factors, setting a new standard for controllable video generation.

视频生成3D控制扩散模型场景编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。