arXiv:2412.11457cs.CV2024-12CVPR被引 11

提升室内多物体新视角合成的结构一致性,避免物体错位与形状不连贯。

MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes

  • 引入深度图和物体掩码增强模型对物体空间关系的理解。
  • 新增物体掩码预测任务,提高物体区分与定位能力。
  • 设计结构引导采样调度器,平衡全局布局与细节恢复。

复用预训练扩散模型在新视角合成(NVS)中已被证明有效,但现有方法多局限于单物体场景;直接应用于多物体组合场景时,常出现物体位置错误、形状与外观不一致等问题,且跨视角一致性缺乏系统评估。为此,我们提出MOVIS,通过改进模型输入、辅助任务和训练策略,增强视图条件扩散模型对多物体场景的结构感知。首先,在去噪U-Net中注入深度图和物体掩码等结构特征,强化模型对物体实例及其空间关系的理解。其次,引入一个辅助任务,要求模型同时预测新视角下的物体掩码,进一步提升物体区分与放置能力。最后,深入分析扩散采样过程,设计结构引导的采样调度器,平衡全局物体布局与细粒度细节恢复的学习。为系统评估合成图像的真实性,我们提出结合跨视角一致性与新视角物体定位评估指标,超越传统图像级指标。在具有挑战性的合成与真实数据集上进行大量实验表明,该方法具备强泛化能力,生成结果具有一致性,展现出对未来3D感知多物体NVS任务的重要潜力。

原文摘要 · Abstract (English)

Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yields inferior results, especially incorrect object placement and inconsistent shape and appearance under novel views. How to enhance and systematically evaluate the cross-view consistency of such models remains under-explored. To address this issue, we propose MOVIS to enhance the structural awareness of the view-conditioned diffusion model for multi-object NVS in terms of model inputs, auxiliary tasks, and training strategy. First, we inject structure-aware features, including depth and object mask, into the denoising U-Net to enhance the model's comprehension of object instances and their spatial relationships. Second, we introduce an auxiliary task requiring the model to simultaneously predict novel view object masks, further improving the model's capability in differentiating and placing objects. Finally, we conduct an in-depth analysis of the diffusion sampling process and carefully devise a structure-guided timestep sampling scheduler during training, which balances the learning of global object placement and fine-grained detail recovery. To systematically evaluate the plausibility of synthesized images, we propose to assess cross-view consistency and novel view object placement alongside existing image-level NVS metrics. Extensive experiments on challenging synthetic and realistic datasets demonstrate that our method exhibits strong generalization capabilities and produces consistent novel view synthesis, highlighting its potential to guide future 3D-aware multi-object NVS tasks. Our project page is available at https://jason-aplp.github.io/MOVIS/.

多物体合成扩散模型新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。