arXiv:2606.00299cs.CVcs.AI2026-06被引 1

用生成的3D缓存增强视频扩散模型,实现精准镜头与物体运动控制。

Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion

论文配图:Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
图 1 · 摘自论文原文
  • 通过3D提升模型构建可编辑的显式3D缓存,提供完整空间先验。
  • 在大视角变化和严重遮挡下保持优异时空一致性,优于纯扩散先验。
  • 适合需要精确控制相机轨迹和多物体运动的视频生成任务。

尽管视频扩散模型(VDMs)在生成高质量视频方面表现优异,但实现精确的相机与场景控制仍具挑战。现有方法主要依赖隐式扩散先验生成未观测区域,导致高动态运动或复杂遮挡时结构崩溃。为此,我们提出Real2SAM2Real,利用3D提升模型(如SAM3D)提取显式可编辑的3D缓存,作为VDM的稳健几何骨架。该缓存捕捉前景实体的完整3D体积而非仅可见外壳,向VDM注入全局空间先验,为复杂场景动态提供可靠的3D感知引导。为有效利用此3D引导同时保留预训练先验,我们设计了软空间对齐注入机制及针对VDM的轻量微调策略。此外,采用掩码法线图作为跨模态桥梁,构建无需3D数据的清洗与扰动流程。大量实验表明,Real2SAM2Real实现了相机轨迹与多实体运动的精确解耦控制。通过引入生成式3D缓存的互补上下文,该框架克服了过度依赖扩散先验导致的典型失效,在大幅相机移动与严重遮挡下仍保持卓越时空一致性。关键在于,通过解耦几何与外观,其专用于VDM的3D缓存消除了因结构空洞、错误立面及反射折射误导带来的视角歧义。

原文摘要 · Abstract (English)

While Video Diffusion Models (VDMs) excel at synthesizing high-fidelity videos, enabling precise camera and scene control remains challenging. Existing methods predominantly rely on implicit diffusion priors to generate unobserved regions, inevitably leading to structural collapse during high-dynamic movements or complex occlusions. To address this challenge, we propose Real2SAM2Real, a framework that leverages 3D lifting models (e.g., SAM3D) to extract an explicitly editable 3D cache, serving as a robust geometric scaffold for the VDM. By capturing the entire 3D volume of foreground entities rather than just their visible shells, this cache injects holistic spatial priors into the VDM, providing dependable 3D-aware guidance for complex scene dynamics. To effectively leverage this 3D guidance while preserving pre-trained priors, we design a Soft Spatial-Aligned Injection mechanism alongside a minimally invasive fine-tuning strategy tailored for VDMs. Furthermore, we employ masked normal maps as a cross-modal bridge to construct a 3D-free data curation and perturbation pipeline. Extensive experiments demonstrate that Real2SAM2Real enables precise, decoupled control over both camera trajectories and multi-entity motions. By utilizing the complementary context from generative 3D caches, our framework overcomes typical breakdowns caused by over-reliance on diffusion priors, maintaining exceptional spatiotemporal consistency under large camera shifts and severe occlusions. Crucially, by decoupling geometry from appearance, our VDM-tailored 3D cache eradicates perspective ambiguities caused by structural holes and erroneous facades, as well as misleading cues from reflections and refractions. Project website is available at https://jiayi-wu-leo.github.io/real2sam2real

视频生成扩散模型3D先验相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。