只需少量关键帧,就能精准插入视频对象并保持真实动态。
PISCO: Precise Video Instance Insertion with Sparse Control
- 用稀疏关键帧控制,自动传播物体外观、运动和场景交互
- 在稀疏条件下的生成稳定性提升,比基线模型更少失真
- 适合影视后期制作,实现高精度、低负担的视频编辑
AI视频生成正从通用生成转向精细可控生成与高保真后处理。在专业影视制作中,精确的局部修改至关重要。核心挑战是视频实例插入:将特定对象精准插入现有画面,同时保持场景完整性。这需要精确时空定位、物理一致的互动关系以及原动态的忠实保留,且用户操作尽可能少。本文提出PISCO,一种基于视频扩散模型的精确视频实例插入方法,支持任意稀疏关键帧控制(单帧、起止帧或任意时间点)。通过可变信息引导、分布保持的时间掩码和几何感知条件建模,解决预训练模型在稀疏条件下的分布偏移问题。构建了包含真实标注与配对干净背景视频的PISCO-Bench基准,采用有参考与无参考感知评估指标。实验表明,PISCO在稀疏控制下持续优于强基线,在增加控制信号时性能单调提升。
原文摘要 · Abstract (English)
The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engineering and "cherry-picking" - towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, it is crucial to perform precise, targeted modifications. A cornerstone of this transition is video instance insertion, which requires inserting a specific instance into existing footage while maintaining scene integrity. Unlike traditional video editing, this task demands several requirements: precise spatial-temporal placement, physically consistent scene interaction, and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To address the severe distribution shift induced by sparse conditioning in pretrained video diffusion models, we introduce Variable-Information Guidance for robust conditioning and Distribution-Preserving Temporal Masking to stabilize temporal generation, together with geometry-aware conditioning for realistic scene adaptation. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms strong inpainting and video editing baselines under sparse control, and exhibits clear, monotonic performance improvements as additional control signals are provided. Project page: xiangbogaobarry.github.io/PISCO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。