arXiv:2512.05672cs.CVcs.AI2025-12被引 7

无需微调视频模型,用隐空间反演实现高效新视角视频生成

InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem

  • 将视频生成转为隐空间的补全反问题,避免微调与标注数据
  • 仅需轻量级掩码编码器,推理时无额外开销且保持生成一致性
  • 保留原始模型能力,适用于通用视频修复与编辑

可控新视角视频生成通常依赖微调预训练视频扩散模型(VDM),但该方法计算成本高且易导致模型原有生成先验的灾难性遗忘。为此,本文提出 InverseCrafter,一种无需训练 VDM 的框架,将新视角视频生成重构为隐空间中的基于补全的反问题,无需任何4D标注训练数据。核心是通过轻量级隐空间掩码编码器,以连续多通道表示建立操作等价性,精准建模隐空间前向过程,实现无需反向传播的高效求解器,跳过重复的VAE运算瓶颈。InverseCrafter在近零额外推理开销下实现高保真、时空一致的新视角合成,并完全保留预训练VDM的生成能力,显著提升通用视频补全与编辑性能。

原文摘要 · Abstract (English)

Recent approaches in controllable novel view video generation often rely on fine-tuning pre-trained Video Diffusion Models (VDMs). This dominant paradigm is computationally expensive and frequently suffers from catastrophic forgetting of the model's original generative priors. To address this challenge, here we propose InverseCrafter, a VDM training-free framework that reformulates novel view video generation as an inpainting-based inverse problem in the latent space, eliminating the need for any annotated 4D training data. The core of our method is to establish operator equivalence by employing a lightweight latent mask encoder to define a latent-domain masking operation via a continuous, multi-channel representation. This principled representation faithfully models the forward process in the latent domain, enabling efficient, backpropagation-free solvers while bypassing the costly bottleneck of repeated VAE operations. InverseCrafter achieves high-fidelity, spatio-temporally coherent novel view synthesis with near-zero additional inference overhead and excels at general-purpose video inpainting and editing by fully preserving the pre-trained VDM's generative capabilities.

视频生成扩散模型隐空间逆问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。