arXiv:2606.01399cs.CV2026-06被引 6

让视频背景随主角动作自然变化,实现真实光影融合。

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion

论文配图:PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
图 1 · 摘自论文原文
  • 用双向注意力联合建模前景动作与背景参考信息
  • 在30万张影视级数据上训练,显著提升动态背景一致性
  • 适合电影制作、虚拟拍摄等需要高保真合成的场景

我们提出PAI-Studio,一种新的参考条件视频生成任务,解决电影级背景替换中的长期难题:在保持前景身份一致、匹配参考场景外观的前提下,生成与前景动作动态对齐且全局光照一致的逼真背景。现有开源系统和商业API难以同时保证运动一致性、高保真前景重光照和身份保留,常导致背景静止、边界不一致及明显合成痕迹。为此,我们基于扩散变换器视频主干,将问题重构为上下文条件生成任务。通过双向注意力机制,模型在统一架构中联合捕捉前景动态与背景参考信息。我们还构建了一个规模达30K的高质量影视与网络视频数据集以支持该任务。大量实验表明,该方法显著优于现有开源及商用API方案。

原文摘要 · Abstract (English)

We present PAI-Studio, a new reference-conditioned video synthesis task that addresses a long-standing challenge in cinematic background replacement: generating dynamic backgrounds aligned with foreground motion while preserving foreground identity, matching reference scene appearance, and achieving globally consistent illumination with realistic foreground relighting. Existing open-source systems and commercial APIs cannot simultaneously ensure motion-consistent background generation, high-fidelity foreground relighting and foreground identity preservation, often resulting in static backgrounds, inconsistent boundaries, and noticeable compositing artifacts. To bridge this gap, we build upon a Diffusion Transformer video backbone and reformulate the problem as an in-context conditional generation task. Through bidirectional attention, our model jointly captures foreground dynamics and background reference information within a unified architecture. We further construct a 30K-scale dataset sourced from high-quality films and online videos to support this task. Extensive evaluations demonstrate that our method significantly outperforms existing open-source and commercial API solutions.

视频生成背景替换扩散模型光影一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。