用2D扩散模型指导多视角图像编辑,保持视图间一致性。
InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization
- 将2D扩散模型的编辑能力迁移到多视角模型,提升跨视角一致性。
- 相比现有方法,视图间一致性显著提升,单帧编辑质量高。
- 适合需要多角度一致编辑的应用,如3D内容生成与虚拟场景设计。
我们解决从稀疏视角输入进行多视角图像编辑的问题,输入可视为从不同视角捕捉场景的图像混合。目标是在遵循文本指令修改场景的同时,保持所有视角的一致性。现有基于神经场或时序注意力的方法在此任务中表现不佳,常产生伪影和不连贯的编辑结果。我们提出InstructMix2Mix(I-Mix2Mix),通过将2D扩散模型的编辑能力蒸馏到预训练的多视角扩散模型中,利用其数据驱动的3D先验实现跨视角一致性。关键贡献是将分数蒸馏采样(SDS)中的传统神经场整合器替换为多视角扩散学生模型,并引入三项新机制:时间步增量式学生更新、专用教师噪声调度以防止退化,以及无需额外开销的注意力改进以增强跨视角连贯性。实验表明,I-Mix2Mix在保持高单帧编辑质量的同时,显著提升了多视角一致性。
原文摘要 · Abstract (English)
We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while preserving consistency across all views. Existing methods, based on per-scene neural fields or temporal attention mechanisms, struggle in this setting, often producing artifacts and incoherent edits. We propose InstructMix2Mix (I-Mix2Mix), a framework that distills the editing capabilities of a 2D diffusion model into a pretrained multi-view diffusion model, leveraging its data-driven 3D prior for cross-view consistency. A key contribution is replacing the conventional neural field consolidator in Score Distillation Sampling (SDS) with a multi-view diffusion student, which requires novel adaptations: incremental student updates across timesteps, a specialized teacher noise scheduler to prevent degeneration, and an attention modification that enhances cross-view coherence without additional cost. Experiments demonstrate that I-Mix2Mix significantly improves multi-view consistency while maintaining high per-frame edit quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。