让角色与环境真实互动,实现电影级视频合成。
Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models

- 通过三掩码引导架构,统一建模角色与环境的双向交互。
- 在真实影视场景中,合成视频质量显著优于现有方法。
- 适合需要精确道具替换和可控环境生成的影视制作团队。
电影级合成旨在将绿幕角色无缝融入新环境,同时保持物理与光照真实感。以往方法难以捕捉角色与周围环境之间的复杂双向交互,我们将其归纳为角色到环境(C2E)的物理作用与环境到角色(E2C)的光照协调。为此,我们提出一个端到端视频扩散框架,联合建模 C2E 与 E2C 交互,特别处理互动道具的挑战。方法引入三掩码引导架构与 RGB-D 联合去噪,确保角色、道具与环境间物理一致性。我们还设计了一种高效的先验驱动数据整理流程,无需昂贵渲染即可构建高质量重光照样本对。此外,参考条件机制支持可控环境生成与精准道具替换。大量实验表明,该框架在动态视频合成质量上显著超越现有方法。
原文摘要 · Abstract (English)
Cinematic compositing aims to integrate green-screen characters into novel environments while maintaining physical and photometric realism. Previous methods often fail to capture the complex bidirectional interactions between characters and their surroundings, which we characterize as Character-to-Environment (C2E) physical interaction and Environment-to-Character (E2C) lighting harmonization. To address this, we propose an end-to-end video diffusion framework that jointly models C2E and E2C interactions, specifically handling the challenges of interactive props. Our approach introduces a tri-mask-guided architecture with RGB-D joint denoising to ensure physically consistent interactions among the character, props, and environment. We further develop an efficient prior-driven data curation pipeline to construct high-quality relighting pairs without expensive rendering. Finally, a reference-conditioned mechanism enables controllable environment synthesis and precise prop replacement. Extensive experiments demonstrate that our framework significantly outperforms existing methods in cinematic-quality dynamic video compositing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。