arXiv:2506.04590cs.CV2025-06被引 65

用视频修复思路实现单目视频的4D生成与编辑,效果更连贯。

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting

  • 将4D生成转为视频修复任务,利用深度点云生成遮挡掩码。
  • 通过自迭代训练提升大视角运动下的时间一致性,生成质量更高。
  • 支持提示词编辑,灵活性强,优于当前最优方法。

我们提出Follow-Your-Creation,一种新型4D视频生成框架,可从单个单目视频输入中同时实现4D内容的生成与编辑。通过利用强大的视频修复基础模型作为生成先验,我们将4D视频生成重新定义为视频修复任务,使模型能够填补因相机轨迹变化或用户编辑导致的缺失内容。为此,我们生成复合掩码修复视频数据,以有效微调模型用于4D视频生成。给定输入视频及其相机轨迹,首先基于深度点云渲染获取不可见区域掩码,指示需补全的区域;同时引入编辑掩码表示用户定义的修改,并将其与不可见掩码结合形成复合掩码数据集。训练时,随机采样不同类型的掩码,构建多样且具挑战性的修复场景,增强模型在各种4D编辑与生成任务中的泛化能力。为应对大相机运动下的时间一致性问题,设计了自迭代调优策略,在每轮微调后使用模型生成下一阶段训练数据,逐步增加视角变化。此外,推理时引入时间打包模块以提升生成质量。该方法有效利用基础模型的先验知识,未损害其原始性能,实现了多视角一致的4D视频生成。同时支持提示词驱动的内容编辑,展现出强大灵活性,在质量和多样性上显著超越现有最先进方法。

原文摘要 · Abstract (English)

We introduce Follow-Your-Creation, a novel 4D video creation framework capable of both generating and editing 4D content from a single monocular video input. By leveraging a powerful video inpainting foundation model as a generative prior, we reformulate 4D video creation as a video inpainting task, enabling the model to fill in missing content caused by camera trajectory changes or user edits. To facilitate this, we generate composite masked inpainting video data to effectively fine-tune the model for 4D video generation. Given an input video and its associated camera trajectory, we first perform depth-based point cloud rendering to obtain invisibility masks that indicate the regions that should be completed. Simultaneously, editing masks are introduced to specify user-defined modifications, and these are combined with the invisibility masks to create a composite masks dataset. During training, we randomly sample different types of masks to construct diverse and challenging inpainting scenarios, enhancing the model's generalization and robustness in various 4D editing and generation tasks. To handle temporal consistency under large camera motion, we design a self-iterative tuning strategy that gradually increases the viewing angles during training, where the model is used to generate the next-stage training data after each fine-tuning iteration. Moreover, we introduce a temporal packaging module during inference to enhance generation quality. Our method effectively leverages the prior knowledge of the base model without degrading its original performance, enabling the generation of 4D videos with consistent multi-view coherence. In addition, our approach supports prompt-based content editing, demonstrating strong flexibility and significantly outperforming state-of-the-art methods in both quality and versatility.

4D生成视频修复单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。