arXiv:2605.12119cs.CVcs.GR2026-05被引 2

用分阶段去噪统一几何与外观生成,解决视图合成中的对齐难题。

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

论文配图:MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
图 1 · 摘自论文原文
  • 早期用几何先验锚定粗结构,后期切换到外观先验修正误差
  • 在点云严重缺失时仍保持鲁棒性,显著优于现有方法
  • 适合需要高精度几何与视觉一致性的3D内容生成场景

生成式新视角合成面临根本矛盾:几何先验能提供空间对齐但随视角变化变得稀疏不准确,外观先验保证视觉保真度却缺乏几何对应。现有方法要么在生成过程中传播几何误差,要么静态融合时产生信号冲突。本文提出MoCam,通过结构化去噪动态,在扩散过程中协调地实现从几何到外观的渐进式演进。模型在早期利用几何先验锚定粗略结构,并容忍其不完整性;后期切换至外观先验,主动修正几何误差并细化细节。该设计自然统一了静态与动态视图合成,通过时间解耦几何对齐与外观精修。实验表明,当点云存在严重孔洞或畸变时,MoCam显著优于先前方法,实现了稳健的几何-外观解耦。

原文摘要 · Abstract (English)

Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors offer visual fidelity but lack geometric correspondence. Existing methods either propagate geometric errors throughout generation or suffer from signal conflicts when fusing both statically. We introduce MoCam, which employs structured denoising dynamics to orchestrate a coordinated progression from geometry to appearance within the diffusion process. MoCam first leverages geometric priors in early stages to anchor coarse structures and tolerate their incompleteness, then switches to appearance priors in later stages to actively correct geometric errors and refine details. This design naturally unifies static and dynamic view synthesis by temporally decoupling geometric alignment and appearance refinement within the diffusion process. Experiments demonstrate that MoCam significantly outperforms prior methods, particularly when point clouds contain severe holes or distortions, achieving robust geometry-appearance disentanglement.

新视角合成扩散模型几何对齐3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。