让视频背景随主角动作自然变化,实现真实光影融合。
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion

- 用双向注意力联合建模前景动作与背景参考信息
- 在30万张影视级数据上训练,显著提升动态背景一致性
- 适合电影制作、虚拟拍摄等需要高保真合成的场景
我们提出PAI-Studio,一种新的参考条件视频生成任务,解决电影级背景替换中的长期难题:在保持前景身份一致、匹配参考场景外观的前提下,生成与前景动作动态对齐且全局光照一致的逼真背景。现有开源系统和商业API难以同时保证运动一致性、高保真前景重光照和身份保留,常导致背景静止、边界不一致及明显合成痕迹。为此,我们基于扩散变换器视频主干,将问题重构为上下文条件生成任务。通过双向注意力机制,模型在统一架构中联合捕捉前景动态与背景参考信息。我们还构建了一个规模达30K的高质量影视与网络视频数据集以支持该任务。大量实验表明,该方法显著优于现有开源及商用API方案。
原文摘要 · Abstract (English)
We present PAI-Studio, a new reference-conditioned video synthesis task that addresses a long-standing challenge in cinematic background replacement: generating dynamic backgrounds aligned with foreground motion while preserving foreground identity, matching reference scene appearance, and achieving globally consistent illumination with realistic foreground relighting. Existing open-source systems and commercial APIs cannot simultaneously ensure motion-consistent background generation, high-fidelity foreground relighting and foreground identity preservation, often resulting in static backgrounds, inconsistent boundaries, and noticeable compositing artifacts. To bridge this gap, we build upon a Diffusion Transformer video backbone and reformulate the problem as an in-context conditional generation task. Through bidirectional attention, our model jointly captures foreground dynamics and background reference information within a unified architecture. We further construct a 30K-scale dataset sourced from high-quality films and online videos to support this task. Extensive evaluations demonstrate that our method significantly outperforms existing open-source and commercial API solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。