arXiv:2605.17423cs.CV2026-05被引 2

用多智能体协作重制长篇电影,保持剧情和角色一致

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

论文配图:Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
图 1 · 摘自论文原文
  • 通过双桥一致性机制,用剧本和视觉锚点维持长期语义连贯
  • 批量生成关键帧并动态验证,减少角色漂移和背景变异
  • 适合影视重制、动画续作等需要长期一致性的场景

我们研究系列级电影重制,即在数百个镜头中进行长时序视频到视频生成,通过风格化或演员替换重制完整剧集或影片,同时严格保持叙事结构、动作编排与角色身份。现有视频生成与编辑流程在此类任务中常因身份漂移、背景变化和语义退化而失效,尤其在大范围摄像机运动与视角变换下。我们提出 Soap2Soap,一种多智能体框架,通过双桥一致性机制实现长期语言-视觉一致性:以场景感知的JSON剧本作为持久语义骨架,并在场景与镜头层级动态分配视觉参考锚点。为抑制合成前的漂移,引入批处理关键帧一致性,通过网格化联合生成多个关键帧于共享潜在空间。闭环验证智能体进一步审计身份、稳定性与对齐性,触发选择性重生成。在 SoapBench 上的实验表明,其在长期一致性与叙事保真度上显著优于商用视频生成API。

原文摘要 · Abstract (English)

We study series-level cinematic remaking, a long-horizon video-to-video generation problem that localizes full episodes or films via stylization or actor replacement while strictly preserving narrative structure, motion choreography, and character identity across hundreds of shots. Existing video generation and editing pipelines often break down in this regime due to compounding identity drift, background mutation, and semantic erosion under large camera motions and viewpoint changes. We propose Soap2Soap, a multi-agent framework that enforces long-term language-visual consistency through a Dual-Bridge Consistency mechanism: a scene-aware JSON screenplay serving as a persistent semantic backbone, and dynamically allocated visual reference anchors at both scene and shot levels. To suppress drift before video synthesis, we introduce batch keyframe consistency, jointly generating multiple keyframes in a shared latent context via a grid-based formulation. A closed-loop verification agent further audits identity, stability, and alignment to trigger selective regeneration. Experiments on SoapBench demonstrate strong improvements over commercial video generation APIs in long-term consistency and narrative fidelity.

视频重制多智能体一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。