首个可精确控制视频内在属性的编辑框架,支持真实感视频生成与修改。
V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties
- 通过逆渲染将视频分解为固有属性通道,实现物理基础的编辑控制。
- 基于关键帧编辑能跨序列保持光照一致性,生成真实感动态视频。
- 适合影视特效、虚拟场景重光照等需要精准控制视觉属性的场景。
大规模视频生成模型在模拟真实场景中的逼真外观与光照交互方面展现出巨大潜力。然而,尚未有闭环框架能同时理解场景固有属性(如反照率、法线、材质和辐照度),利用这些属性进行视频合成,并支持可编辑的固有表示。我们提出 V-RGBX,首个端到端的固有感知视频编辑框架。V-RGBX 集成三大能力:(1) 视频逆渲染生成固有通道,(2) 从这些固有表示中合成逼真视频,(3) 基于关键帧的固有通道条件化视频编辑。核心是交错条件机制,使用户通过选择关键帧即可直观地操控任意固有属性。大量定性和定量实验表明,V-RGBX 能生成时间一致、逼真的视频,并以物理合理方式传播关键帧编辑。我们在物体外观编辑和场景级重新照明等应用中验证了其有效性,性能超越已有方法。
原文摘要 · Abstract (English)
Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsic scene properties (e.g., albedo, normal, material, and irradiance), leverages them for video synthesis, and supports editable intrinsic representations remains unexplored. We present V-RGBX, the first end-to-end framework for intrinsic-aware video editing. V-RGBX unifies three key capabilities: (1) video inverse rendering into intrinsic channels, (2) photorealistic video synthesis from these intrinsic representations, and (3) keyframe-based video editing conditioned on intrinsic channels. At the core of V-RGBX is an interleaved conditioning mechanism that enables intuitive, physically grounded video editing through user-selected keyframes, supporting flexible manipulation of any intrinsic modality. Extensive qualitative and quantitative results show that V-RGBX produces temporally consistent, photorealistic videos while propagating keyframe edits across sequences in a physically plausible manner. We demonstrate its effectiveness in diverse applications, including object appearance editing and scene-level relighting, surpassing the performance of prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。