arXiv:2607.29627cs.CV2026-07

让静态图和动态视频都能精准按轨迹合成,不依赖3D重建。

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

论文配图:FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
图 1 · 摘自论文原文
  • 用统一表征分离物体运动与位置,适配不同输入
  • 无需参数就能把物体沿指定路径精准放置
  • 适合影视特效、虚拟拍摄等需要高精度合成的场景

生成式视频合成需将外部素材无缝融入现有视频序列,但现有方法存在控制与保真度的权衡:或从静态图幻化运动,无法保留预动画素材动态;或缺乏精细空间控制,难以按用户定义轨迹精准放置。本文提出FlexComposer,将视频合成统一为轨迹引导的条件生成任务,实现静态图与动态视频的无缝融合。核心设计包括:(1) 统一的本体表征,解耦物体内在运动与整体位移,将异构输入标准化至稳定居中隐空间;(2) 空间感知隐空间注入,利用VAE隐空间的平移等变性,通过无参数机制将本体特征投射至目标轨迹;(3) 混合数据集与仿真到真实课程,结合程序化模拟、真实电影画面与生成数据,隐式学习物理合理的光照与阴影融合。该统一设计可处理产品图到动态主体等多种输入,实现高保真运动控制与环境融合,无需显式3D重建或可学习适配器。大量实验表明,FlexComposer在视觉质量、时序一致性和轨迹遵循上优于现有最先进方法。

原文摘要 · Abstract (English)

Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.

视频合成轨迹控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。