用电影级动态场景替换多个角色,保持画面一致性与视觉质量。
PAI-Actor: Cinematic Multi-Character Replacement in Dynamic Scenes

- 基于结构引导的人体重建思路,实现多角色动画生成
- 1080P高清短片段生成,支持长视频高效推理
- 适合影视制作、虚拟演员替换等实际应用
我们提出PAI-Actor,一个用于动态电影场景中多角色替换的电影级动画框架。不同于传统仅驱动单个静态图像或单一主体的系统,我们的目标是在真实视频片段中替换并动画化多个角色,同时保留原始场景动态、摄像机运动和背景内容。该任务极具挑战性,因为生成角色必须在动作和交互上与源表演一致,并在光照、阴影、构图和整体电影质感上匹配周围环境。为此,我们将多角色动画建模为结构引导的人体恢复问题,并构建了基于高质量电影数据的电影驱动训练流程。为进一步支持实际影视生产,我们引入双向到自回归的蒸馏框架:先训练一个双向扩散Transformer以在1080P分辨率下生成高质量短片段,再将其蒸馏为自回归视频到视频模型,实现高效推理与长视频生成。实验表明,PAI-Actor可实现高保真多角色动画,具备强场景一致性、电影级视觉质量和高效的长时生成能力。
原文摘要 · Abstract (English)
We present PAI-Actor, a cinematic multi-character animation framework for character replacement in dynamic movie scenes. Unlike conventional animation systems that mainly drive a single static image or a single subject, our goal is to replace and animate multiple characters within real video clips while preserving the original scene dynamics, camera motion, and background content. This setting is particularly challenging because the generated characters must remain consistent with the source performance in motion and interaction, while also matching the surrounding background in lighting, shadow, composition, and overall cinematic appearance. To address this, we formulate multi-character animation as a structure-guided human recovery problem and build a movie-driven training pipeline from high-quality film data. Furthermore, to support practical cinematic production, we introduce a bidirectional-to-autoregressive distillation framework: we first train a bidirectional diffusion transformer for high-quality short-clip generation at 1080P resolution, and then distill it into an autoregressive video-to-video model for efficient inference and longer video generation. Experiments show that PAI-Actor enables high-fidelity multi-character animation with strong scene consistency, cinematic visual quality, and efficient long-form generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。